AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas sharing a 256 GiB LMCache pool. After adopting an existing packed-page patch, we isolate a raw-pointer fallback that omits the dependency on the current CUDA stream. Controlled byte tests fail under an imposed delay and pass when the dependency is restored; the existing mixed allocator provides a working deployment path. Full-pool allocation checks and service regression c

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.