llama.cpp/ggml/src at b7960 - llama.cpp - Gitea: Git with a cup of tea

sdgoij/llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-05-08 10:04:10 +00:00

Files

History

Jeff Bolz 1946e46f4c vulkan: For coopmat2 FA, use fp16 accumulators for the final result (#19376 )

The cpu and cuda backends use fp16 for the VKQ accumulator type, this change
does the same for vulkan. This helps particularly with large head sizes which
are very register-limited.

I tried this for the coopmat1 path and it slowed down a bit. I didn't try for
scalar.

I applied the softmax bias that the cuda backend uses to avoid overflow,
although I was not able to reproduce the original bug without it.

2026-02-06 09:15:13 +01:00

..

ggml : add ggml_build_forward_select (#18550 )

2026-01-19 20:03:19 +02:00

docs : Minor cleanups (#19252 )

2026-02-02 08:38:55 +02:00

ggml-cpu: use LUT for converting e8->f32 scales on x86 (#19288 )

2026-02-04 09:43:29 +08:00

cuda : cuda graphs now compare all node params (#19383 )

2026-02-06 07:55:06 +02:00

ggml-hexagon: flash-attention and reduce-sum optimizations (#19141 )

2026-01-30 21:14:20 -08:00

HIP: add mmf for CDNA (#18896 )

2026-01-29 11:10:53 +01:00

metal : skip loading all-zero mask (#19337 )

2026-02-06 09:25:11 +02:00

CUDA: faster tile FA, add oob checks, more HSs (#16492 )

2025-10-11 20:54:32 +02:00

opencl: refactor some ops, concat, repeat, tanh and scale (#19226 )

2026-02-02 15:54:43 -08:00

rpc : use unordered_map::reserve and emplace (#18513 )

2026-01-02 12:09:36 +02:00

Remove support for Nvidia & AMD GPU, because the oneAPI plugin for Nvidia & AMD GPU is unavailable: download/installation channels are out of work. (#19246 )

2026-02-02 21:06:21 +08:00

ggml-virtgpu: make the code thread safe (#19204 )

2026-02-04 10:46:18 +08:00

vulkan: For coopmat2 FA, use fp16 accumulators for the final result (#19376 )

2026-02-06 09:15:13 +01:00

Remove pipeline cache mutexes (#19195 )

2026-02-01 18:47:29 -08:00

ggml-zdnn : mark zDNN buffers as non-host (#18967 )

2026-01-22 01:16:21 +01:00

ggml-zendnn : resolve ZenDNN backend cross-module symbol dependency (#19159 )

2026-01-29 12:28:57 +08:00

CMakeLists.txt

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-alloc.c

llama: automatically set parameters not set by the user in such a way that maximizes GPU utilization (#16653 )

2025-12-15 09:24:59 +01:00

ggml-backend-dl.cpp

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend-dl.h

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend-impl.h

llama: use host memory if device reports 0 memory (#18587 )

2026-01-09 05:34:56 +08:00

ggml-backend-reg.cpp

hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )

2026-01-29 12:33:21 -08:00

ggml-backend.cpp

ggml-backend: fix async set/get fallback sync (#19179 )

2026-02-02 10:00:05 +01:00

ggml-common.h

llama : add gpt-oss (#15091 )

2025-08-05 22:10:36 +03:00

ggml-impl.h

ggml : add ggml_build_forward_select (#18550 )

2026-01-19 20:03:19 +02:00

ggml-opt.cpp

finetune: SGD optimizer, more CLI args (#13873 )

2025-08-14 12:03:57 +02:00

ggml-quants.c

ggml : fix uninitialized is_on_grid in quantize_row_iq3_xxs_impl (#15928 )

2025-09-23 10:25:20 +02:00

ggml-quants.h

llama : add gpt-oss (#15091 )

2025-08-05 22:10:36 +03:00

ggml-threading.cpp

ggml : build backends as libraries (#10256 )

2024-11-14 18:04:35 +01:00

ggml-threading.h

remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797 )

2024-12-12 19:02:49 +01:00

ggml.c

ggml: added cleanups in ggml_quantize_free (#19278 )

2026-02-03 08:43:39 +02:00

ggml.cpp

ggml : Print backtrace on uncaught C++ exceptions (ggml/1232)

2025-06-01 13:43:57 +03:00

gguf.cpp

GGUF: check that tensor size is representable (#19072 )

2026-01-24 21:57:51 +01:00