CVE-2026-43632

Summary

llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when –sleep-idle-seconds is configured.

Affected Software

VendorProductVersion RangeStatus
ggml-orgllama.cppb7492 <= b9060affected

Weaknesses

  • CWE-416: CWE-416 Use After Free
  • CWE-367: CWE-367 Time-of-check Time-of-use (TOCTOU) Race Condition

References