Tools · · 2 min read

llama.cpp b11425 fixes a CUDA Qwen3.6 prompt-processing regression that llama-bench missed

llama.cpp b11425 changes CUDA graph optimisation so allocation dependencies do not vary with micro-batch size. It fixes a reported Qwen3.6-35B-A3B prompt-processing regression in a real completion workload, not a new serving feature.


ggml-org published the pre-release llama.cpp b11425 on October 5, 2026. Its sole listed change is “CUDA: make the alloc_deps check batch independent,” delivered in PR #29986. (Source: llama.cpp b11425 release notes, 2026-10-05)

Key facts:

  • b11425 is a pre-release published on October 5, 2026.
  • The change is limited to the CUDA backend.
  • The upstream issue used Qwen3.6-35B-A3B-GGUF on three NVIDIA RTX PRO 6000 Blackwell GPUs.
  • In that reporter’s completion workload, prompt throughput fell from 13,816.18 to 6,858.24 tokens/s after an earlier change.
  • The same reporter’s llama-bench result stayed near 15,600 tokens/s on both builds.
Official GitHub release page for llama.cpp b11425 showing its CUDA allocation-dependency change.
b11425 is an upstream nightly pre-release with one named CUDA change. Screenshot: ggml-org release notes.

What this means if you serve Qwen with llama-server

This is a useful reminder not to promote a synthetic benchmark into a production readiness check. The linked bug report says a Qwen3.6-35B-A3B CUDA deployment regressed after PR #29184. Its author found no material difference with llama-bench at a prompt length of 18,606 tokens, but saw approximately half the prior prompt-evaluation speed with llama-completion and the same regression in llama-server. (Source: llama.cpp issue #29980, 2026-10-05)

The reported good build processed the prompt at 13,816.18 tokens/s. The reported bad build processed it at 6,858.24 tokens/s. Those are one user’s measurements on a three-GPU Blackwell machine, not a general performance promise for every Qwen quantization or CUDA host.

PR #29986 moves the batch-size-sensitive MMVQ fusion check out of CUDA graph optimisation. The code comment explains why: graph optimisation must keep the same topology for every micro-batch size, otherwise ggml-alloc has to reserve again and the scheduler has to synchronize during runtime. (Source: PR #29986 patch, merged 2026-10-05)

Official llama.cpp GitHub issue documenting the Qwen3.6 prompt processing performance regression on CUDA.
The original report distinguishes an unchanged synthetic benchmark from a slower real prompt-evaluation path. Screenshot: llama.cpp issue #29980.

If a CUDA-backed Qwen server recently became slow on long repository or retrieval prompts, test b11425 in staging. Capture time to first token and prompt-evaluation time with a representative request, not just llama-bench. The upstream build guide documents -DGGML_CUDA=ON for a CUDA build. (Source: llama.cpp CUDA build documentation, retrieved 2026-10-07)

This release does not establish a gain for Apple Metal, Vulkan, CPU inference, or decode speed. For repeat-prompt workloads, combine that staging check with the llama-server KV-cache reuse guide. If several agents share the endpoint, the llama-server concurrency guide helps isolate slot pressure from prompt-processing regressions.

Sources

Source: ggml-org/llama.cpp b11425 release notes