llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-05-11 19:44:06 +00:00

Files

Devedse fd89556567 [SYCL] Add BF16 support to GET_ROWS operation (#21391 )

Add GGML_TYPE_BF16 to the SYCL backend's GET_ROWS operation, both in
supports_op and in the kernel dispatch. This fixes a performance
regression where models using BF16 embedding tensors (e.g., Gemma4's
per_layer_token_embd.weight) fall back to CPU for the GET_ROWS op,
causing a full GPU-to-CPU tensor transfer every token.

The fix reuses the existing get_rows_sycl_float template with
sycl::ext::oneapi::bfloat16, matching the pattern already used for
sycl::half (F16) and float (F32).

2026-05-09 08:50:24 +03:00

cmake

ggml: backend-agnostic tensor parallelism (experimental) (#19378 )

2026-04-09 16:42:19 +02:00

include

CUDA: lower-case PCI bus id, standardize for ggml (#22820 )

2026-05-08 10:09:38 +02:00

src

[SYCL] Add BF16 support to GET_ROWS operation (#21391 )

2026-05-09 08:50:24 +03:00

.gitignore

vulkan : cmake integration (#8119 )

2024-07-13 18:12:39 +02:00

CMakeLists.txt

ggml : bump version to 0.11.0 (ggml/1478)

2026-05-05 13:15:59 +03:00