Evidence

Each result keeps its boundary.

A result fixes source, GPU, workload, method, and excluded claims.

Primary results

−61.846%

Mistral KV

4 x 12K prompts with 19,077 physical slots

0.9662×

Median runtime

19,077 PureSWA versus 50,000-slot reference

+25.81%

Full capacity

real gpt-oss-20b, same 1.979 GiB KV budget

−20.30%

Owner vs Stock

balanced four-way real-checkpoint ablation

−42.105%

Head-stripe KV

exact multi-scale-window reduction versus max-window allocation

Evidence

ResultValueContractBoundary
Uniform SWA execution−61.846%Mistral-7B, 4 x 12K prompts, 19,077 slots versus 50,000one H20; page 1; three balanced pairs; identical output digests
Real checkpoint capacity+25.81%openai/gpt-oss-20b, Full capacity 47,616 to 59,904same 1.979 GiB KV budget; no radix/spec/overlap/Graph
Real checkpoint makespan−20.30%Owner32 vs Stock128, four balanced execution orders8 x 6000 prompt + 32 decode; identical output-token digests
Capacity prediction4 / 4predicted Full/SWA pools matched fresh SGLang processes exactlygpt-oss-20b, one H20, recorded 1.979 GiB KV budget
Chunked Local synthesisResettablefloor(q/C) == floor(k/C) to epoch arena and chunk-end deathhost exhaustive proof; SGLang lowering not yet enabled
Lifetime Normalization−42.105%32 KV heads with 512/2048/8192 token windowsexact host geometry and Manager proof; no GPU claim

Exclusions

  • No released-checkpoint model-quality improvement claim.
  • No radix/prefix-cache, speculative, overlap, or CUDA Graph qualification.
  • No claim that VMM already backs SGLang KV tensors.
  • No multi-GPU multicast or fabric-memory claim.
  • One H20 does not qualify other NVIDIA architectures.

The complete result index and source hashes are in results/.