Keeping agent conversations cached
On NInfer, a long coding-agent conversation could lose its cache once the system-memory tier filled, so its next turn re-read 200K tokens from scratch. We fixed how the engine decides where that state fits.
- Published
- Topics
- Speed
- Setup
- RTX 5090 · Qwen3.8-27B NVFP4 · NInfer
113×
faster turn after overflow
99.996%
of the prompt served from cache
Same
speed when nothing overflows
The result
Why a lost cache hurts
A cached turn only processes the new message. A lost cache means recomputing the entire prompt before the first token appears, and that cost grows with the conversation. Agent sessions reach 100K to 200K tokens within an hour of work.
What went wrong
NInfer keeps an active conversation's state on the GPU. When another conversation needs the GPU, that state moves to a tier of pinned system memory (RAM), so the next turn can resume instead of recomputing.
The system-memory tier is one memory arena shared by three allocation sizes: state images of 147 MiB, KV pages of 2.0 MiB and draft-model KV pages of 129 KiB. Each allocation needs one contiguous free range. The engine judged room by counting free bytes. After enough turnover the arena was fragmented, and in one traced failure it had 35.0 MB free against a 33.8 MB move, all of it in gaps smaller than a single KV page.
The engine approved the move without freeing anything, the move found nowhere to go, and the engine then released the conversation it had set out to preserve. From that point the tier kept its old conversations and lost every new one.
The fix
- Room means placement. The arena simulates placing the exact allocations a move needs, after the ranges that evicting a candidate would return.
- Eviction stops when the state fits. The engine evicts older conversations until the incoming state can be placed, prefers an eviction that completes placement over one that only frees bytes, and evicts nothing when no choice would make it fit.
- One rule for every system-memory write. State captures and pause snapshots use the same placement check, and page reservations follow the same grouping as moves.
The tier recycles instead of freezing
Seven 200K-token conversations were sent in turn for 100 rounds, then an eighth arrived. During the rounds both builds served every turn from cache in under 0.5 s. The difference appears once the tier is full.
Why 3 and 6 rather than 1 and 2: the fix evicts the conversations whose memory frees one contiguous range for the new one, so memory layout decides, not age. A block-based layout for this tier (#379) would remove the fragmentation and let the engine choose by recency and value instead, keeping the conversations an agent is most likely to return to.
In two further runs the fix held up across sizes and concurrency. With 16 conversations cycling through 16.5K, 66K, 131K and 200K tokens, every repeat send and 28 of 29 checks on an earlier conversation were served from cache. With two agents working at the same time on a full tier, every repeat send and recheck of the pair was served from cache in 0.23 to 0.64 s.
When it matters
The fix changes nothing while every active conversation fits on the GPU. On an RTX 5090 with this configuration the GPU holds about 254K tokens of KV cache, roughly one long agent session. The system-memory tier, and with it this fix, comes into play when work outgrows that:
- an agent that runs sub-agents, or two agents working at once
- switching between several long sessions or projects
- long-running sessions where hours of turnover fragment the tier
Recommendations
If you run long agent sessions on NInfer
Use a build with the fix: the
saylekbranch of Saylek-ai/ninfer today, or upstream once it merges. Give the system-memory tier as much pinned memory as the machine can spare with--host-context-mib; it decides how many evicted conversations stay resumable.If you maintain NInfer
The change is proposed upstream as three pull requests on #378. The remaining cost, evicting two conversations to admit one, comes from fragmentation itself. #379 proposes a block-based arena that would remove it.
What we measure next
A two-agent replay on a full tier, stock and fixed on the same upstream base, to put a number on the gain for concurrent sessions in addition to the single-overflow case shown here.
Method
| Item | Detail |
|---|---|
| Hardware | One NVIDIA RTX 5090, 50 GiB pinned system-memory tier |
| Model | Qwen3.8-27B NVFP4 (qwen3_8_27b_nvfp4.ninfer), 240K-token context, two lanes, MTP speculative decoding |
| Stock build | NInfer 68c54356 |
| Fixed builds | 68c54356 + fix; 81c8ce09 + fix |
| Common to all | An unrelated tool-call parsing fix (#244) |
| Prompts | Synthetic, seeded random-word conversations sized with the engine's tokenizer; 16-token replies |
| Timing | Client wall clock per request on the same machine |
| Reproduction | Script and engine flags attached to #378 |
All figures come from synthetic runs on a test instance. The runs isolate one failure: the system-memory tier full and a conversation that must move to it. How often that happens in practice depends on how many long sessions share one GPU.