llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-05-11 11:34:10 +00:00

Author	SHA1	Message	Date
Sumit Chatterjee	1e5ad35d56	model : add sarvam_moe architecture support (#20275 ) b9093	2026-05-09 16:31:50 +02:00
Yuannan	65d7a8bbf0	devops : updated Nix systems (#22869 )	2026-05-09 17:15:03 +03:00
Davi Henrique Linhares	00d56b11c3	docker : upgraded the default intel compute-runtime version (#22567 )	2026-05-09 10:22:23 +02:00
Alessandro de Oliveira Faria (A.K.A.CABELO)	5757c4dcb1	cmake : update BoringSSL to 0.20260508.0 (#22839 ) b9090	2026-05-09 10:26:33 +03:00
Alexey Kopytko	e20b83930c	SYCL: reduce allocation overhead during flash attention (#22732 ) * SYCL: reduce allocation overhead during flash attention * tidy up whitespace * add a note about the flag * move ggml_sycl_fattn_* into fattn-buffers.hpp * refactor implementation into fattn-buffers.cpp * move new_fattn_kv_buffers back into ggml-sycl.cpp b9089	2026-05-09 09:30:39 +03:00
Devedse	fd89556567	[SYCL] Add BF16 support to GET_ROWS operation (#21391 ) Add GGML_TYPE_BF16 to the SYCL backend's GET_ROWS operation, both in supports_op and in the kernel dispatch. This fixes a performance regression where models using BF16 embedding tensors (e.g., Gemma4's per_layer_token_embd.weight) fall back to CPU for the GET_ROWS op, causing a full GPU-to-CPU tensor transfer every token. The fix reuses the existing get_rows_sycl_float template with sycl::ext::oneapi::bfloat16, matching the pattern already used for sycl::half (F16) and float (F32). b9088	2026-05-09 08:50:24 +03:00
Intel AI Get-to Market Customer Success and Solutions	60489932ec	sycl: Q5_K reorder MMVQ/dequant + Q8_0 reorder MMVQ path (#22152 ) * sycl: Q5_K reorder MMVQ/dequant + Q8_0 reorder MMVQ path Signed-off-by: Chun Tao <chun.tao@intel.com> * Remove duplicate definitions --------- Signed-off-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Todd Malsbary <todd.malsbary@intel.com> b9087	2026-05-09 08:48:07 +03:00
Intel AI Get-to Market Customer Success and Solutions	4a4f819cb6	sycl: Battlemage AOT build via spir64_gen + MMQ subgroup annotations (#22147 ) * sycl: Battlemage AOT build via spir64_gen + MMQ subgroup annotations Signed-off-by: Chun Tao <chun.tao@intel.com> * Remove unneeded/unnecessary comments and annotations The MMQ subgroup annotations added are on functions gated behind ggml_sycl_supports_mmq(). Revisit the need for these annotations when that function changes. --------- Signed-off-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Todd Malsbary <todd.malsbary@intel.com>	2026-05-09 08:42:40 +03:00
AesSedai	046e284437	Add flash attention MMA / Tiles to support MiMo-V2.5 (#22812 ) * mimo-v2.5: add flash attention mma/tiles for for d_kq=192 d_v=128 * mimo-v2.5: follow (256, 256) fattn templates * mimo-v2.5: cleanup comments * mimo-v2.5: further comment cleanup * mimo-v2.5: address PR feedback fix GQA handling check for other dangling 320/576 carveouts and mirror them for 192 Add to backend ops test so new paths are covered b9085	2026-05-09 11:28:29 +08:00
Yanzhao Wang	66001722aa	hexagon: add HTP kernel for GGML_OP_GATED_DELTA_NET (#22837 ) Implement the Gated Delta Net recurrence on HVX with: - 4-row fused kernels for PP (prompt processing) path - 8-row fused kernels for TG (token generation) path, reducing K/Q/gate vector reload overhead by 2x - Separate PP/TG thread functions for I-cache isolation - VTCM state scratchpad with DMA in/out for TG single-cycle access - Vectorized gate exp via hvx_exp_f32 b9084	2026-05-08 17:12:04 -07:00
Intel AI Get-to Market Customer Success and Solutions	c5703e03a5	sycl: support non-contiguous input in PAD op (#22148 ) Signed-off-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Todd Malsbary <todd.malsbary@intel.com>	2026-05-09 08:05:22 +08:00
Pranav Dhinakar	b46812de78	Feature hexagon l2 norm (#22816 ) * L2_NORM Updates * Addressed PR Comments * ggml-hexagon: add L2_NORM HVX kernel for Hexagon backend * hex-unary: remove supported_unary_nc since the outer loop is the same for all unary ops --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com> b9082	2026-05-08 13:41:40 -07:00
Aldehir Rojas	49956041ee	common : do not wrap raw strings in schema parser for tagged parsers (#22827 ) b9081	2026-05-08 15:33:17 -05:00
ynankani	9f5f0e689c	model : support Gemma4_26B_A4B_NVFP4 (#22804 ) * Gemma4_26B_A4B_NvFp4 hf checkpoint convert to gguf format fixes Signed-off-by: ynankani <ynankani@nvidia.com> * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Address review comments Signed-off-by: ynankani <ynankani@nvidia.com> * fix CRLF Signed-off-by: ynankani <ynankani@nvidia.com> * Lint error fix Signed-off-by: ynankani <ynankani@nvidia.com> --------- Signed-off-by: ynankani <ynankani@nvidia.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9080	2026-05-08 20:42:09 +02:00
Aldehir Rojas	f9cd456ea5	common : revert reasoning budget +inf logit bias (#22740 ) b9079	2026-05-08 17:46:43 +02:00
smugman-dot	5d6f18a638	webui: fix LLM title generation for agentic conversations (#22840 )	2026-05-08 16:36:04 +02:00
Xuan-Son Nguyen	29debb3a6a	server: support Vertex AI compatible API (#22545 ) * server: support Vertex AI compatible API * a bit safer * support other AIP_* env var * various fixes * if AIP_MODE is unset, do nothing * fix test case * fix windows build b9077	2026-05-08 15:23:04 +02:00
Xuan-Son Nguyen	9dcf835528	server: (router) expose child model info from router's /v1/models (#22683 ) * server: (router) expose child model info from router's /v1/models * update docs b9076	2026-05-08 14:42:15 +02:00
Pascal	58e68df0f9	cuda: fuse snake activation (mul, sin, sqr, mul, add) (#22667 ) * cuda: fuse snake activation (mul, sin, sqr, mul, add) Add ggml_cuda_op_snake_fused with F32 / F16 / BF16 templates. The matcher recognizes the naive 5 op decomposition emitted by audio decoders (BigVGAN, Vocos) for snake activation y = x + sin(ax)^2 inv_b and rewrites it to a single elementwise kernel. Add test_snake_fuse comparing CPU naive vs CUDA fused across F32 / F16 / BF16. * cuda: address review feedback from @am17an Use ggml_cuda_cast for F32/F16/BF16 conversions and rename kernel_snake to snake_kernel to match upstream conventions. * cuda: snake fusion fastdiv on T_len, Suggested-by: @am17an * Update tests/test-backend-ops.cpp Co-authored-by: Aman Gupta <amangupta052@gmail.com> * cuda: snake fusion check add->type matches x->type Address review feedback from @am17an * cuda: snake fusion check add->type matches x->type Moved for readability (equivalent) Address review feedback from @am17an --------- Co-authored-by: Aman Gupta <amangupta052@gmail.com> b9075	2026-05-08 17:44:09 +08:00
Aleksander Grygier	9b2925e1e0	webui: Add Import/Export of Settings configuration + improve architecture (#22803 ) * refactor: Settings keys as constant object keys * chore: Run `npm audit fix` * refactor: Settings Sections UI * feat: Refactor Settings structure and implement import/export logic * feat: Introduce ROUTES constant and RouterService * refactor: Consolidate settings definitions into registry * refactor: Update settings page routing structure * chore: Migrate hardcoded URLs to use ROUTES and RouterService * feat: Enhance model selection logic for settings and chat * chore: Update webui static build * refactor: Address PR review comments * fix: Remove unneeded setting * fix: Re-add missing settings * fix: Add missing `/slots` proxy for webui dev mode * chore: Dev-mode logs * fix: Data binding * fix: Steering for non-agentic flow	2026-05-08 11:26:04 +02:00
Johannes Gäßler	a8fd165fec	CUDA: lower-case PCI bus id, standardize for ggml (#22820 ) b9073	2026-05-08 10:09:38 +02:00
miyan	6d57a49a70	vulkan: fix spv shadowing (#22760 ) b9072	2026-05-08 09:35:22 +02:00
Max Krasnyansky	3e941b813b	ggml: update SCHED_DEBUG output to use ggml_op_desc() (#22825 ) b9071	2026-05-07 22:43:04 -07:00
Shawn Gu	f3e8d149ce	opencl: add q4_0 MoE GEMM for Adreno (#22731 ) * Q4_0 MoE CLC pass sanity check * release program * opencl: fix whitespace * opencl: remove unused cl_program * opencl: break #if block to make it more clear * opencl: adjust format --------- Co-authored-by: Li He <lih@qti.qualcomm.com> b9070	2026-05-07 21:17:07 -07:00
Michał Piszczek	1d72d87349	convert : fix RuntimeError when stripping FP8 KV-cache scales (#22818 ) * convert : fix RuntimeError when stripping FP8 KV-cache scales In ModelBase._generate_nvfp4_tensors the final cleanup loop iterates self.model_tensors.keys() and calls del on the same dict, which raises RuntimeError: dictionary changed size during iteration when a ModelOpt NVFP4 model also has FP8 KV-cache scales (e.g. mmangkad/Qwen3.6-35B-A3B-NVFP4 and any modelopt config with kv_cache_quant_algo: FP8). Wrap the keys view in list() so the deletions happen on a snapshot. * re-add another accidentally removed list --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-05-08 06:55:48 +03:00
Neo Zhang	6a2a2513dc	fix script error (#22795sycl : )	2026-05-08 06:54:57 +03:00
samuraieng	44dbe8c521	model: Support sarashina2.2-vision-3b model (#22103 )	2026-05-07 23:10:29 +02:00
leonardHONG	05ff59cb57	CUDA: batch out_prod inner loop with cublasSgemmStridedBatched (#22651 ) * CUDA: batch out_prod inner loop with cublasSgemmStridedBatched * CUDA: batch out_prod inner loop with cublasSgemmStridedBatched * CUDA: add cublasSgemmStridedBatched mapping for HIP and MUSA backends b9066	2026-05-07 21:59:29 +02:00
smugman-dot	aaf4a4d5e0	webui: add option for LLM title generation (#22265 ) * webui: add LLM title generation option * webui: use chat_template_kwargs for title gen + fix conversation check * webui: capture firstUserMessage before async streamChatCompletion to fix race condition * webui: extract LLM title generation into separate method * webui: use constants and ChatService for LLM generated titles * webui: rebuild static output * webui: add LLM title generation setting to new settings location * webui: use sendMessage in generateTitle * webui: rebuild static output * webui: fix formatting * webui: configurable title prompt, remove think tag regexes, fix TS error * webui: group title constants into TITLE object, use TruncatedText for CSS truncation and fix race condition * webui: rebuild static output	2026-05-07 21:14:03 +02:00
Georgi Gerganov	e43431b381	llama : fix device state save/load (#22805 ) b9064	2026-05-07 21:43:40 +03:00
shaofeiqi	ceb7e14b96	opencl: add opfilter regex for debugging (#22782 ) b9063	2026-05-07 11:00:20 -07:00
Aldehir Rojas	093be624cc	common/chat : preserve media markers for typed-content templates (#22634 ) b9062	2026-05-07 12:50:56 -05:00
HaoJun ZHANG	deab41ec68	tests: add long-sequence cases and fix inputs for gated_delta_net (#22794 ) * tests : add long-seq + tail cases for gated_delta_net * tests : realistic input ranges for gated_delta_net b9061	2026-05-08 00:23:36 +08:00
Intel AI Get-to Market Customer Success and Solutions	ad09224658	sycl: add FILL, CUMSUM, DIAG, SOLVE_TRI, SSM_SCAN, GATED_DELTA_NET (#22149 ) * sycl: add FILL, CUMSUM, DIAG, SOLVE_TRI, SSM_SCAN, GATED_DELTA_NET Signed-off-by: Chun Tao <chun.tao@intel.com> * Fix abort during test-backend-ops Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * Regenerate ops.md Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * Add scope_dbg_print to newly added SYCL ops. Also add scope_dbg_print to existing ssm_conv op. Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> --------- Signed-off-by: Chun Tao <chun.tao@intel.com> Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> Co-authored-by: Chun Tao <chun.tao@intel.com> Co-authored-by: Todd Malsbary <todd.malsbary@intel.com> b9060	2026-05-07 18:51:33 +03:00
Gaurav Garg	b9afc19cb4	Write a readme on Multi-GPU usage in llama.cpp (#22729 ) * Write a readme on Multi-GPU usage in llama.cpp * Apply suggestions from code review Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Address review comments * Apply suggestions from code review Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>	2026-05-07 17:48:40 +02:00
Georgi Gerganov	803627f121	llama : remove unnecessary seq_id check during state restore (#22797 ) b9058	2026-05-07 16:37:26 +03:00
pl752	68380ae11b	ggml-cpu: Optimized risc-v cpu q1_0 dot b9057	2026-05-07 21:09:25 +08:00
Pascal	cc97e45a14	mtmd: fix whisper audio tail truncation by exposing padded buffer to FFT (#22770 ) b9056	2026-05-07 14:01:01 +02:00
AesSedai	8e52631d55	model: Add Mimo v2.5 model support (#22493 ) * add mimo-v2.5 support * mimo-v2.5: fix modify_tensors row split * mimi-v2.5: forgot `add_attn_value_scale` plumbing * mimi-v2.5: fix tp dequant to detect tp rows * mimo-v2.5: fix TP iteration to be descending * mimo-v2.5: fix comment * mimo-v2.5: retain fused qkv * mimo-v2.5: missed the attn_value scale during merge * mimo-v2.5: fused QKV needs contiguous for scaling attention value * mimo-v2.5: move `speech_embeddings.` to TextModel filter_tensors * Update src/llama-hparams.h Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update src/models/mimo2.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update src/models/mimo2.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update src/models/mimo2.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * mimo-v2.5: include MTP weights in gguf --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9055	2026-05-07 13:21:58 +02:00
Pascal	f4b5a2ee91	webui: fix ?model= URL param race in router mode (#22771 ) * webui: fix ?model= URL param race in router mode * chore: update webui build output	2026-05-07 13:09:32 +02:00
Vishal Singh	97f06e9eed	codeowners : add ZenDNN backend codeowner (#22772 ) * codeowners : add ZenDNN backend codeowner * codeowners : fix zendnn owners to use individual github handles	2026-05-07 14:46:51 +08:00
viggy	e358d75adb	webui: fix flicker issue on dismiss animation on overlay primitives (#22773 ) * add fill-mode-forwards * generated diffs	2026-05-07 08:11:31 +02:00
Shane Tran Whitmire	cfff1fc300	sycl : fix test script (#22737 ) The error: ./examples/sycl/test.sh: line 122: level_zero:${$GGML_SYCL_DEVICE}: bad substitution was thrown whenever the user used this command: ./examples/sycl/test.sh -mg 0 Fix is to get rid of a dollar sign.	2026-05-07 08:25:57 +03:00
Adrien Gallouët	3980e04d5a	llama : add missing call to ggml_backend_load_all() (#22752 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co> b9050	2026-05-07 08:24:47 +03:00
tc-mb	2496f9c149	mtmd : support MiniCPM-V 4.6 (#22529 ) * Support MiniCPM-V 4.6 in new branch Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix code bug Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix pre-commit Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix convert Signed-off-by: tc-mb <tianchi_cai@icloud.com> * rename clip_graph_minicpmv4_6 Signed-off-by: tc-mb <tianchi_cai@icloud.com> * use new TYPE_MINICPMV4_6 Signed-off-by: tc-mb <tianchi_cai@icloud.com> * use build_attn to allow flash attention support Signed-off-by: tc-mb <tianchi_cai@icloud.com> * no use legacy code, restored here. Signed-off-by: tc-mb <tianchi_cai@icloud.com> * use the existing tensors name Signed-off-by: tc-mb <tianchi_cai@icloud.com> * unused ctx->model.hparams.minicpmv_version Signed-off-by: tc-mb <tianchi_cai@icloud.com> * use n_merge for slice alignment Signed-off-by: tc-mb <tianchi_cai@icloud.com> * borrow wa_layer_indexes for vit_merger insertion point Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix code style Signed-off-by: tc-mb <tianchi_cai@icloud.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * use filter_tensors and add model.vision_tower Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix chkhsh Signed-off-by: tc-mb <tianchi_cai@icloud.com> * fix type check Signed-off-by: tc-mb <tianchi_cai@icloud.com> --------- Signed-off-by: tc-mb <tianchi_cai@icloud.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9049	2026-05-06 21:54:09 +02:00
Gilad S.	5207d120ea	model : don't crash on unsupported architecture (#22742 ) * model: don't crash on unsupported architecture * Update src/llama-model.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> b9048	2026-05-06 18:51:21 +02:00
fl0rianr	a0101225bc	common: do not fit to unknown device memory (#22614 ) * common: do not fit to unknown device memory Signed-off-by: Florian Reinle <f.reinle@otec.de> * common: preserve host fallback for non-GPU fit devices Signed-off-by: Florian Reinle <f.reinle@otec.de> * common: keep unknown GPU fit memory at zero Signed-off-by: Florian Reinle <f.reinle@otec.de> --------- Signed-off-by: Florian Reinle <f.reinle@otec.de> b9047	2026-05-06 17:03:45 +02:00
Georgi Gerganov	a290ce6266	gguf-py : bump version to 0.19.0 (#22664 ) * gguf-py : bump version to 0.19.0 * bump poetry --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> gguf-v0.19.0	2026-05-06 14:46:14 +02:00
Yakine Tahtah	a00e47e422	mtmd: add granite-speech support (ibm-granite/granite-4.0-1b-speech) (#22101 ) * mtmd: add granite-speech support (ibm-granite/granite-4.0-1b-speech) Conformer encoder with Shaw relative position encoding, QFormer projector, log-mel spectrogram with frame stacking. Encoder uses GLU gating, folded batch norm, and SSM depthwise conv. QFormer compresses encoder output via windowed cross-attention (window=15, queries=3) into the LLM embedding space. Audio preprocessing: reflect-padded STFT, 80-bin mel filterbank, dynamic range compression, 2x frame stacking (80->160 mel). GGUF converter handles batch norm folding at export time, fused K/V split, and Conv1d weight reshaping. Tested against HF transformers reference: token-for-token match on 30s/60s audio clips with greedy decoding. * mtmd: rename gs_ prefixed tensors to generic/architecture names * mtmd: use tensor_mapping.py for all granite_speech tensors * convert: fold GraniteSpeechTextModel into GraniteModel * mtmd: replace n_layer hack with explicit has_standard_layers flag * mtmd: replace hardcoded magic numbers with GGUF hparams for granite speech * mtmd: align KEY_A_ define spacing * convert: register GraniteModel for GraniteSpeechForConditionalGeneration * convert: fix ty type-check for GraniteSpeechMmprojModel registration * mtmd: align TN_ define spacing * mtmd: use generic layer loop for granite speech tensor loading * mtmd: merge qformer_proj_layer into clip_layer * mtmd: granite_speech remove redundant ggml_build_forward_expand on inputs * mtmd: granite_speech add comment explaining why build_attn is not used * mtmd: granite_speech hard-code eps in cpp, remove from GGUF metadata * gguf: add spacing between granite_speech tensor mapping blocks * mtmd: make generic audio layer_norm_eps read optional * mtmd: granite_speech keep encoder eps in GGUF, only hard-code projector eps * mtmd: align defines and struct fields in clip-impl.h and clip-model.h * mtmd: fix alignment and ordering issues across granite speech files * convert: granite_speech use filter_tensors instead of modify_tensors for skipping b9045	2026-05-06 14:40:59 +02:00
David Huggins-Daines	750141969c	feat: migrate to PEP 621 and add uv support (#21907 ) * feat: migrate to PEP 621 and add uv support * fix: remove upper bound on protobuf * remove poetry.lock and uv.lock * fix/add torch dependency version and markers * fix dev-dependency deprecation warning * gguf-py : update python version requirement to 3.10 --------- Co-authored-by: David Huggins-Daines <dhd@dhd.ecolingui.ca> Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2026-05-06 14:04:10 +02:00

1 2 3 4 5 ...

9093 Commits