Primary results
Mistral KV
4 x 12K prompts with 19,077 physical slots
Median runtime
19,077 PureSWA versus 50,000-slot reference
Full capacity
real gpt-oss-20b, same 1.979 GiB KV budget
Owner vs Stock
balanced four-way real-checkpoint ablation
Head-stripe KV
exact multi-scale-window reduction versus max-window allocation
Evidence
| Result | Value | Contract | Boundary |
|---|---|---|---|
| Uniform SWA execution | −61.846% | Mistral-7B, 4 x 12K prompts, 19,077 slots versus 50,000 | one H20; page 1; three balanced pairs; identical output digests |
| Real checkpoint capacity | +25.81% | openai/gpt-oss-20b, Full capacity 47,616 to 59,904 | same 1.979 GiB KV budget; no radix/spec/overlap/Graph |
| Real checkpoint makespan | −20.30% | Owner32 vs Stock128, four balanced execution orders | 8 x 6000 prompt + 32 decode; identical output-token digests |
| Capacity prediction | 4 / 4 | predicted Full/SWA pools matched fresh SGLang processes exactly | gpt-oss-20b, one H20, recorded 1.979 GiB KV budget |
| Chunked Local synthesis | Resettable | floor(q/C) == floor(k/C) to epoch arena and chunk-end death | host exhaustive proof; SGLang lowering not yet enabled |
| Lifetime Normalization | −42.105% | 32 KV heads with 512/2048/8192 token windows | exact host geometry and Manager proof; no GPU claim |
Exclusions
- No released-checkpoint model-quality improvement claim.
- No radix/prefix-cache, speculative, overlap, or CUDA Graph qualification.
- No claim that VMM already backs SGLang KV tensors.
- No multi-GPU multicast or fabric-memory claim.
- One H20 does not qualify other NVIDIA architectures.
The complete result index and source hashes are in results/.