llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-05-15 13:34:06 +00:00

Author	SHA1	Message	Date
Xuan Son Nguyen	457fbdac2c	fix compile	2025-11-21 23:26:32 +01:00
Xuan Son Nguyen	525e2746df	address review comments	2025-11-21 23:25:34 +01:00
Xuan Son Nguyen	b0540e8e1e	add env for args	2025-11-21 23:06:49 +01:00
Xuan Son Nguyen	7241558835	better --models-dir	2025-11-21 23:06:09 +01:00
Xuan Son Nguyen	7cd929076d	remove default model path	2025-11-21 22:33:04 +01:00
Xuan Son Nguyen	62ee883d5a	implement LRU	2025-11-21 22:22:57 +01:00
Xuan Son Nguyen	032b9ff4a9	add --models-dir param	2025-11-21 11:11:01 +01:00
Xuan Son Nguyen	a2e912cf35	address review comment	2025-11-20 21:54:22 +01:00
Xuan Son Nguyen	cd5c699304	add docs (first version)	2025-11-20 21:45:05 +01:00
Xuan Son Nguyen	be25bccdff	address review comment	2025-11-20 21:37:22 +01:00
Xuan Son Nguyen	6929c9f43d	address thread safety issue	2025-11-20 18:38:02 +01:00
Xuan Son Nguyen	5369aaa1d6	address most problems	2025-11-20 18:34:22 +01:00
Xuan Son Nguyen	216140867e	tmp apply upstream fix	2025-11-20 18:19:21 +01:00
Xuan Son Nguyen	5805ca7960	add is_active()	2025-11-20 16:26:31 +01:00
Xuan Son Nguyen	d0ea9e0830	also allow terminate loading model	2025-11-20 16:20:14 +01:00
Xuan Son Nguyen	6610724f8e	fix unsafe pointer	2025-11-20 16:13:30 +01:00
Xuan Son Nguyen	b9ebdf616a	more stable	2025-11-20 15:49:40 +01:00
Xuan Son Nguyen	919d3f8cbf	Merge branch 'master' into xsn/server_model_management_v1_2	2025-11-20 14:19:16 +01:00
Aleksander Grygier	4c91f2633f	Improved file naming & structure for UI components (#17405 ) * refactor: Component iles naming & structure * chore: update webui build output * refactor: Dialog titles + components namig * chore: update webui build output * refactor: Imports * chore: update webui build output	2025-11-20 14:07:31 +01:00
Piotr Wilkin (ilintar)	92c0b387a9	grammar : fix integer overflow (#17381 ) * Fix DoS / integer overflow * Remove optional, use INT64_MAX instead as placeholder value (it's technically -1, so it fits :) * White space * Actually, since it's unsigned, use UINT64_MAX b7118	2025-11-20 14:47:04 +02:00
Xuan Son Nguyen	7c6eb17fad	fix windows	2025-11-20 13:14:56 +01:00
Georgi Gerganov	2286a360ff	sync : ggml b7117	2025-11-20 14:10:44 +02:00
YangLe	1d321e592b	metal : fix compile on macos 11 (whisper/3533)	2025-11-20 14:10:44 +02:00
Georgi Gerganov	196f5083ef	common : more accurate sampling timing (#17382 ) * common : more accurate sampling timing * eval-callback : minor fixes * cont : add time_meas impl * cont : fix log msg [no ci] * cont : fix multiple definitions of time_meas * llama-cli : exclude chat template init from time measurement * cont : print percentage of unaccounted time * cont : do not reset timings	2025-11-20 13:40:10 +02:00
o7si	5088b435d4	convert : fix TypeError when loading base model remotely in convert_lora_to_gguf (#17385 ) * fix: TypeError when loading base model remotely in convert_lora_to_gguf * refactor: simplify base model loading using cache_dir from HuggingFace * Update convert_lora_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * feat: add remote_hf_model_id to trigger lazy mode in LoRA converter --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-11-20 12:30:12 +01:00
Piotr Wilkin (ilintar)	845f200b28	ggml : Fix transposed SOLVE_TRI result (#17323 ) * Did someone transpose the SOLVE_TRI result matrix? Perhaps... * Update ggml/src/ggml-cpu/ops.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update ggml/src/ggml-cpu/ops.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b7113	2025-11-20 12:58:21 +02:00
Scott Fudally	a7784a8b1d	DGX Spark: UMA support (#17368 ) * DGX Spark: UMA support * Updates from PR feedback * More PR feedback cleanup * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * Remove trailing whitespace * Update ggml/src/ggml-cuda/ggml-cuda.cu --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> b7112	2025-11-20 12:32:02 +02:00
Adrien Gallouët	79bb743512	ggml : remove useless and error-prone variadic macros (#17399 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co> b7111	2025-11-20 11:18:27 +01:00
sudhiarm	3ae282a06f	kleidiai: fix zero-size array declaration (#17240 ) b7110	2025-11-20 11:45:49 +02:00
ixgbe	5be353ec4a	ggml-cpu:add RISC-V RVV (Zvfh) optimization for FP16 vector scaling (#17314 ) * ggml-cpu:add RISC-V RVV (Zvfh) optimization for FP16 vector scaling Signed-off-by: Wang Yang <yangwang@iscas.ac.cn> * fix comment * fix comment 2 --------- Signed-off-by: Wang Yang <yangwang@iscas.ac.cn> b7109	2025-11-20 08:09:18 +02:00
Xuan Son Nguyen	0ef3b61e82	add test	2025-11-20 00:29:59 +01:00
Xuan Son Nguyen	5423d42a35	use subprocess.h, better logging	2025-11-20 00:05:29 +01:00
Xuan Son Nguyen	54b3545791	fix windows build	2025-11-19 22:30:47 +01:00
Xuan Son Nguyen	abc0ca478a	does this fix windows?	2025-11-19 22:24:00 +01:00
Xuan Son Nguyen	399f536dc7	fix compile error	2025-11-19 21:33:44 +01:00
Xuan Son Nguyen	fc5901a449	server: add model management and proxy	2025-11-19 21:23:00 +01:00
Giuseppe Scrivano	7d77f07325	vulkan: implement ADD1, ARANGE, FILL, SOFTPLUS, STEP, ROUND, CEIL, FLOOR, TRUNC (#17319 ) * vulkan: initialize array * vulkan: implement ADD1 * vulkan: implement ARANGE * vulkan: implement FILL * vulkan: implement SOFTPLUS * vulkan: implement STEP * vulkan: implement ROUND * vulkan: implement CEIL * vulkan: implement FLOOR * vulkan: implement TRUNC * docs: update Vulkan ops Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> b7108	2025-11-19 17:29:45 +01:00
Jeff Bolz	1fa4551af0	vulkan: support larger argsort (#17313 ) * vulkan: support larger argsort This is an extension of the original bitonic sorting shader that puts the temporary values in global memory and when more than 1024 threads are needed it runs multiple workgroups and synchronizes through a pipelinebarrier. To improve the memory access pattern, a copy of the float value is kept with the index value. I've applied this same change to the original shared memory version of the shader, which is still used when ncols <= 1024. * Reduce the number of shader variants. Use smaller workgroups when doing a single pass, for a modest perf boost * reduce loop overhead * run multiple cols per invocation, to reduce barrier overhead b7107	2025-11-19 17:25:50 +01:00
Jeff Bolz	2eba631b81	vulkan: Add copy_transpose shader (#17371 ) b7106	2025-11-19 16:50:43 +01:00
Aleksander Grygier	99c53d6558	webui: Add a "Continue" Action for Assistant Message (#16971 ) * feat: Add "Continue" action for assistant messages * feat: Continuation logic & prompt improvements * chore: update webui build output * feat: Improve logic for continuing the assistant message * chore: update webui build output * chore: Linting * chore: update webui build output * fix: Remove synthetic prompt logic, use the prefill feature by sending the conversation payload ending with assistant message * chore: update webui build output * feat: Enable "Continue" button based on config & non-reasoning model type * chore: update webui build output * chore: Update packages with `npm audit fix` * fix: Remove redundant error * chore: update webui build output * chore: Update `.gitignore` * fix: Add missing change * feat: Add auto-resizing for Edit Assistant/User Message textareas * chore: update webui build output	2025-11-19 14:39:50 +01:00
Sigbjørn Skjæret	07b0e7a5ac	convert : use self.block_count everywhere instead of reading hparams (#17359 )	2025-11-19 11:52:38 +01:00
Aman Gupta	fd7353d5eb	cuda: fix rope fusion for gemma3 (#17378 ) b7103	2025-11-19 18:25:05 +08:00
Piotr Wilkin (ilintar)	6fd4f95367	Fix too relaxed check on CUDA "fast copy" (can_be_transposed) condition (#17332 ) * Fix too relaxed check on CUDA "fast copy" (can_be_transposed) condition * Argh. * Making CISC happy ;) * Integrate CONT tests * Use loopy loop * Skip new tests for (B)F16 for now. b7102	2025-11-19 10:36:33 +01:00
Ruben Ortlam	980b7cd17e	vulkan: force full subgroups for flash attention to fix intel subgroup crash (#17356 ) b7101	2025-11-19 08:46:26 +01:00
Jeremy Rand	c49daff5ba	ggml-cpu: Don't pass -mpowerpc64 when -mcpu already implies it (#17308 ) b7100	2025-11-19 14:19:00 +08:00
Xuan-Son Nguyen	10e9780154	chat: fix int overflow, prevent size calculation in float/double (#17357 ) * chat: fix int overflow, prevent size calculation in float/double * Update common/chat.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> b7099	2025-11-18 19:11:53 +01:00
Haiyue Wang	a045492088	vocab : call reserve() for building plamo-2-translate suffix (#17343 ) Test 'Q4_K_M' quantization on https://huggingface.co/pfnet/plamo-2-translate The 'suffix_to_score' size is 193510, it needs 19 memory allocation with final capacity 262144 to hold the value, if not preserve the memory. Signed-off-by: Haiyue Wang <haiyuewa@163.com>	2025-11-18 18:58:22 +01:00
hksdpc255	1920345c3b	common : Generalized XML-style tool-call parsing with streaming support (GLM 4.5/4.6 + MiniMax M2 + SeedOSS + Kimi-K2 + Qwen3-Coder + Apriel-1.5 + Xiaomi-MiMo) (#16932 ) * Add files via upload * fix unit test * fix crashes for --reasoning-format=none * Patch buggy official MiniMax-M2 chat template * add upstream minja fix: https://github.com/ochafik/minja/pull/7 * Fix <think> token not generated * add test copied from https://github.com/ggml-org/llama.cpp/pull/16946 * cleanup * Hopes to fix the compilation error on CI * Delete chat template patching since it’s fixed by upstream Minja * Remove undeeded Minimax-M2 template patch https://github.com/ochafik/minja/pull/7#issuecomment-3480356100 * Add proper handling of optional parameters with test merged tests from: `23d4bb75c4` * Fix making all tool parameters optional * Move xml tool parser to separate file * cleanup & add tests for GLM4.5 * add streaming tests & enhancement & cleanups Add streaming test for both GLM 4.5 and minimax-m2. Cleanup for preserved_tokens. Cleanup for grammar rule name. Enhance the parser's stability. * cleanup & add support for Kimi-K2 Qwen3-Coder Apriel-1.5 Xiaomi-MiMo * apply suggestions from reviewers * fix a misuse for data.grammar_lazy * fix grammar when tool have no argument * Fix `no triggers set for lazy grammar!` for GLM4.5/4.6. Insert additional stops for Kimi-K2 * update chat.cpp * fix grammar for GLM 4.5/4.6 * Try fix Jinja template for GLM * Try fix GLM-4.6.jinja * Update common/chat-parser-xml-toolcall.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update tests/test-chat.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * improve chat template for GLM, rename Kimi-K2 template to Kimi-K2-Thinking * Improve Kimi-K2 chat template * Fix unit test * Fix "Invalid tool call arguments passed" in a rare case. In a rare case, the model may emit a raw string that begins with a valid JSON string. This commit adds unit tests to cover that scenario and fixes the regression introduced during the Kimi-K2 adaptation. --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b7097	2025-11-18 18:54:15 +01:00
jiahao su	561a3e2788	ci : change the openEuler-310p image to fix release (#17361 ) b7096	2025-11-18 18:10:23 +01:00
Georgi Gerganov	f40a2e5f11	gitignore : be more specific about ignored stuff (#17354 )	2025-11-18 16:44:53 +02:00

1 2 3 4 5 ...

7144 Commits