Researchers propose KVBoost for faster LLM inference via chunk-level KV cache reuse
KVBoost reduces time-to-first-token by 4.49x in evaluations on Qwen2.5-3B by enabling cache reuse beyond contiguous prefixes, with no measured accuracy loss.
1 source · cross-referenced
- KVBoost introduces a chunk-level KV cache reuse system for Hugging Face-compatible decoder models, enabling reuse regardless of where shared content appears in prompts.
Transformer-based large language models incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions.
KVBoost, a chunk-level KV cache reuse system for Hugging Face-compatible decoder models, enables reuse regardless of content position by introducing a dual-hash keying scheme that separates positional identity (prefix hash) from content identity (content hash), supporting both exact and approximate cache matches.
To address attention boundary errors from independently cached chunks, KVBoost employs two repair strategies: SelectiveRecompute, which re-encodes boundary regions, and CacheBlendRecompute, which identifies and recomputes high-deviation tokens after a probe pass.
The system further incorporates asymmetric KV quantization (int8/int4), adaptive chunk boundary splitting, and importance-weighted eviction under a fixed memory budget.
Evaluated on Qwen/Qwen2.5-3B over 1,000 bug-localization samples, KVBoost achieves a 4.49x reduction in time-to-first-token (142.4 ms vs. 639.1 ms) and outperforms prefix caching by 16%, with no loss in accuracy (99.2% vs. 99.1%).
KVBoost provides a practical, memory-bounded inference acceleration layer compatible with RoPE-based models without architectural modification.
- Aug 25, 2026 · arXiv cs.AI
Paper surveys risks of synthetic-data training loops and mitigation strategies
Trust79 - Aug 25, 2026 · arXiv cs.CL
Paper proposes two distinct update operators for incremental narrative interpretation in AI systems
Trust79 - Aug 24, 2026 · arXiv cs.CL
Paper quantifies ‘clinical lost-in-the-middle’ effect in EHR processing and proposes query-conditioned context selection
Trust79