llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-05-12 12:04:08 +00:00

Files

Max Krasnyansky f5d1c4179f hexagon: dma optimizations (mostly fixing regressions) (#21137 )

* hex-fa: add simple dma cache for Mask

I noticed that we were refetch the mask rows over and over.
This simple cache avoids that.

* hex-dma: unset in-order desc bit which caused signficant perf regression

We don't rely on true in order processing of the DMA descriptors anywhere.
Turns out this mode caused significant regression of around 3-4 TPS during token gen.

* hex-rope: update comment to clarify that we don't need in-order DMA completions

2026-03-29 06:40:13 -07:00

cmake

ggml: Skip backend library linking code when GGML_BACKEND_DL=ON (#15094 )

2025-08-07 13:45:41 +02:00

include

llama: fix llama-model-saver (#20503 )

2026-03-25 12:53:16 +02:00

src

hexagon: dma optimizations (mostly fixing regressions) (#21137 )

2026-03-29 06:40:13 -07:00

.gitignore

vulkan : cmake integration (#8119 )

2024-07-13 18:12:39 +02:00

CMakeLists.txt

ggml : bump version to 0.9.8 (ggml/1442)

2026-03-18 15:17:28 +02:00