DocumentationArchitecture

OrbitKV architecture

Mission

OrbitKV is a KV cache for vLLM and SGLang and a proposed framework-neutral state planner. It does not schedule model execution. Each framework owns its HBM allocation and active GPU page lifecycle. Its adapter exposes block identity and registered GPU buffers; OrbitKV currently owns external pinned DRAM/SSD replicas and transfer leases. SGLang and vLLM hybrid layouts share compiled page demand and recovery validation; general lifetime analysis, retention and joint placement/routing policy remain future work.

The vLLM connector and SGLang direct GPU linker share the same Rust cache client and have passed single-node GPU recovery tests.

Process topology

Current compiled page demand, engine ownership and cache tiers; future lifetime and physical planning

Run one independent OrbitKV Cache Manager per inference host, with one or more engine instances connected to it. Framework adapters run in the inference processes and use the same cache API for local DRAM, SSD, and remote fetches. The cache manager decides where to source a hit; the inference engine still decides when to query and save. Remote fetch is experimental. There is no OrbitKV KV-aware request router today.

       current multi-node cache (experimental)
   host A                                      host B
   vLLM or SGLang                             vLLM or SGLang
   engine-owned HBM                           engine-owned HBM
        | CUDA IPC + UDS/iceoryx2                   | CUDA IPC + UDS/iceoryx2
   Cache Manager A ---- Mooncake RDMA/TCP ---- Cache Manager B
   pinned DRAM / SSD                         pinned DRAM / SSD
            \                                   /
             \---- etcd members/placement ----/
      catalog shards embedded in Managers; one copy per shard

Single-node deployment connects engines to their host’s Cache Manager and needs neither Catalog nor peer gRPC. The Manager shares external capacity across instances; model/storage identities still determine whether bytes are reusable. Container GPU/PID/IPC wiring and concurrent multi-engine serving require separate qualification. Current SSD backing is a cache file truncated on Cache Manager startup, not durable KV storage across manager restarts. Distributed Managers advertise sealed replicas to assigned catalog shards using cached membership, then query missing evidence in bounded batches. They authorize/pin source data before Mooncake reads bytes. Catalog restart is repaired from surviving owner inventories. Each shard has one metadata copy; replication and online placement handoff remain future work.

Standalone deployment has no gRPC listener. Registration, health, sessions, and cleanup use the authenticated bootstrap UDS. --etcd-endpoints with Node ID and catalog placement enables a peer gRPC listener for catalog synchronization, discovery, source authorization and lock release. Process IPC supports query, publish, asynchronous restore completion, and lease release: iceoryx2 carries fixed descriptors while a Unix socket authenticates the peer, passes a sealed memfd descriptor arena, and supplies an eventfd for wakeups. The vLLM adapter requires this path and fails fast if the Cache Manager socket is missing. Each inference process must reach a Cache Manager on its own host. Pending queries return Loading and continue on Tokio. The endpoint owns one session-scoped operation/revision registry bound to instance, request, and group; the core query future owns its backing reads and returns a terminal result. Cancelling or disconnecting drops reply ownership while submitted reads drain, including cache admission and lease release, without another poll. SGLang’s plugin admission hook keeps pending requests queued until a leased result or bounded fallback is available. vLLM defers further lookup admission until an admitted restore reaches its first compute step. This prevents a deferred lookup that cannot allocate GPU pages from stranding a completed restore behind it in the waiting queue. Query reservations use the registered group’s padded bytes and remain charged through preparation, result ownership, and GPU completion. Global and instance limits bound retained payloads; identical backing reads can be shared while each request keeps its own ticket and lease. See query budgets. Users configure SSD paths and capacity; engine adapters do not choose the storage backend, and normal deployment leaves --ssd-backend at its default. With automatic SSD selection, the Manager tries native cuFile on ext4/XFS and falls back to io_uring when unavailable. On the cuFile path a demand result can own a pinned file extent instead of host bytes. A dedicated GPU storage worker reads through bounded registered staging and scatters only the selected state. DRAM restores, and speculative preparation retain their existing paths. Complete groups can be written from GPU staging; fragmented groups seal in DRAM before writeback. GPU-storage files reserve physical capacity before admission. Reads coalesce across source leases within each file, while the restore task retains all leases and excludes unrequested aligned gaps. Two registered 4 MiB slots issue asynchronous cuFile I/O with event/byte-count completion checks. The storage queue limits GPU writes to eight jobs and one in-flight write, rotates jobs by batch and bounds read bursts; saturation uses host publication/io_uring. Native GDS qualification is separate from compatibility-mode correctness. Publish holds its iceoryx2 reply until D2H and any GPU-backed SSD writes finish, so the caller does not release source HBM pages early while the dispatcher remains free. The Rust cache client opens a separate descriptor session for Publish on its first save, so an in-flight save does not serialize the worker’s Query/Restore calls behind that reply. Instance cleanup serializes against registration, drains GPU load/save/storage queues, and only then releases imported CUDA mappings. Superseded sessions cannot clean up a replacement session. Both vLLM and the SGLang direct linker register CUDA IPC pages and use iceoryx2 descriptors on the hot path. Remote transfers use the Mooncake-backed TransferEngine. See transport.md for the measured process-transport baseline.

API and crate boundaries

LayerCodeOwns
Framework adapterspython/orbitkv/vllm, python/orbitkv/sglangFramework-specific hashes, layout, and page-lifetime events
Cache clientorbitkv-channel/src/cache_client.rs, python/src/client.rsRust query/warming ownership, independent publish session, client-bound restore handles and GIL-free waiting; PyO3 API
Connection setuppython/orbitkv/client/connection.pyEngine endpoint options and same-host socket selection
State contractorbitkv-stateState identity, format compatibility, compiled page demand, recovery validation, page-reference types
Process IPCorbitkv-channel, orbitkv-server/src/endpoint/iceoryx2 requests/replies, UDS bootstrap and lifecycle, pending queries, descriptor generation
Process utilitiesorbitkv-commonShared logging setup and peer connection defaults
Hardware localityorbitkv-core/src/memory/numa.rsNUMA topology and allocation/worker affinity
Cache statisticsorbitkv-server/src/metric/hll.rsNamespaced miss cardinality and windowed reuse estimates
Cache serviceorbitkv-server/src/cache/Transport-neutral operations, registration, and session cleanup
Cache engineorbitkv-coreLeases, HBM transfer scheduling, pinned DRAM, SSD, local and remote lookup
Peer controlorbitkv-server/src/peer.rs, orbitkv-core/src/peer/export.rsServer translates RPCs; Core validates and owns source grants
Replica catalogorbitkv-catalog, orbitkv-core/src/peer/catalogCandidate ownership and node liveness; embedded fixed shards with cached member admission
Byte movementorbitkv-transfer, orbitkv-mooncake-sysMooncake Segment/BatchTransfer over RDMA or TCP

Transport-specific names belong at physical boundaries. Cache operations and framework adapters use placement-neutral names and results. Moving a cache hit from DRAM to SSD or another node should not change query_prefetch, save, start_restore, or release for the caller. The process channel implements the current iceoryx2/UDS connection without defining a separate cache API.

Core module ownership

ModuleResponsibility
engine/Instance registration, EngineConfig, Publish orchestration, demand validation and restore handoff
memory/NUMA placement, pinned allocations and pools
storage/Residency assembly and allocator-driven reclamation; publish.rs owns queued sealing and publication
storage/dram/Resident images, eviction/admission policy, exact insertion versions and inventory
storage/ssd/Files, index, immutable extent leases, io_uring/cuFile I/O and registered staging
planning/Metadata-only discovery, batch replica evidence, completion targets and source/path eligibility
query/Admission budgets, shared reads, host materialization, query phases and leases
peer/Catalog client, cached candidates, authoritative exports, requester READs and completion recovery
transfer/Registered engine layouts, GPU copies/codecs and completion-drained workers
codec/Representation validation and encoding/decoding
cost/Operation observations, bounded estimates and same-target shadow comparisons

lib.rs defines the public API. Tests mirror these modules under crates/orbitkv-core/tests/unit/; GPU integration gates stay in tests/. backing/ and internode/ have been removed. There is one SSD store with independent access routes; peer transport is not a storage medium. PeerExports checks live owner/version evidence and holds source memory until completion. The Mooncake registration owner retains its pinned pool through unregister. Cost observations and shadow comparisons remain opt-in and do not select a new execution route. Remote SSD/HBM and GPU-direct cache endpoints remain future work.

Restore returns one completion receiver after all submitted DMA drains. The old shared-memory completion state and its second load API have been removed.

Upstream designs and OrbitKV owners

LMCache, FlexKV and Mooncake provide implementation references for concrete cache mechanisms. OrbitKV applies them through its existing state contract and Rust resource owners. vLLM and SGLang adapters continue to supply engine layouts, scheduler signals and page ownership; they do not gain separate cache schedulers. The complete implementation plan also maps LMCache MP deployment, prefetch/store policies, lazy offload, allocation/event sharing and instance isolation to concrete OrbitKV work.

ReferenceMechanism to useOrbitKV owner and status
LMCache v0.5.5 GDS contextPreallocated storage, reusable registered staging, stream-ordered I/O with retained submission statestorage/ssd reserves capacity; cufile/slot owns registered streams/staging and stable asynchronous arguments/results through event completion.
LMCache MP serialization and FlexKV compressionSeparate engine precision from cache encoding; bound codec workspace and qualify formatscodec/ owns batched GPU ANS/FP8/TurboQuant, reusable arenas, CPU SIMD and CRC validation; transfer/worker/codec owns engine-page and writeback lifetimes. Encoded DRAM, SSD and Mooncake payloads share versioned metadata. cuFile can write encoded GPU groups and restore through GPU validation/decode. Native GDS and broader model-quality qualification remain open.
FlexKV file-range coalescing and GDSMerge physically compatible same-file ranges; keep storage geometry separate from engine tensor layoutstransfer/worker/ssd validates demand and coalesces leased ranges per file; its queue owns task/extent lifetime, bounded GPU write admission and batch-level read/write scheduling.
Mooncake TE v0.3.13.post1Registered memory and batched remote transfersReused directly through orbitkv-transfer and orbitkv-mooncake-sys. Catalog/source authorization and state compatibility remain OrbitKV responsibilities. Two-host/RDMA qualification is still pending.
Mooncake RFC #3504 — draft proposalCached membership and embedded authority; keep coordination off per-key data pathsorbitkv-catalog and server/cluster already use embedded shards, cached membership and etcd leases/Watch. Catalog replication, online placement and repair remain future work; the RFC is not evidence that those features are implemented.

Compiled required_ranges, complete-state recovery and generation/lease checks remain the common acceptance boundary for every tier. A useful transfer policy must reduce request latency or resource cost under matched workloads without weakening those checks. Reuse, prefetch and retention policies keep their existing evidence gates; an upstream default alone does not justify enabling an OrbitKV policy. The roadmap keeps single-node correctness, DP sharing, P/D reuse and catalog availability separately qualified.

Layering

vLLM adapter                SGLang adapter
block hashes / CUDA IPC     radix hashes / CUDA IPC
                             /
       PyO3 CacheManagerClient (cache API)
                    |
       Rust CacheClient (request ownership)
                    |
    orbitkv-channel / iceoryx2 + UDS
                    |
              orbitkv-server/cache/operations
                           |
                    orbitkv-core
                 cache · leases · tiers
                    /           \
                  SSD      Mooncake Transfer
                           |
                    peer DRAM / SSD

     peer control: tonic / gRPC, only with distributed etcd/placement configuration

    orbitkv-state: shared state identity and recovery semantics

orbitkv-state

This crate contains no framework or CUDA dependencies. Its first public types are:

  • StateKey: the materialized model/storage namespace and versioned native prefix/group key;
  • StateDescriptor: logical token span, component and format evidence for future recovery validation;
  • StateFormat: model/implementation digest, dtype, layout, and parallel shape;
  • StateComponent: attention KV, MLA, recurrent, convolution, SWA, draft, and indexer state;
  • StateBundle, RecoveryContract and RecoveryDemand: recovery evidence, compiled rules and the complete selected boundary’s required group ranges.

RecoveryContract::compile normalizes declared prefix/window/checkpoint rules once at registration. required_ranges(namespace, start, end) exposes absolute page-aligned intervals per group from the engine’s valid HBM prefix origin. Prefix demand covers the tail, window demand rounds up to pages and is capped by that tail, and checkpoint demand selects its final page. restorable_boundaries uses the same requirements to check namespace identity, aligned spans and complete leased coverage. Both adapters intersect legal boundary sets across ranks or shards; hybrid reconciliation is shared.

SGLang checks exact transferred-plus-retained keys against those ranges; vLLM uses them for hybrid allocation. Hybrid discovery returns metadata-only candidate positions; Rust computes legal boundaries and read_recovery uses those same ranges to slice actual reads and revalidate leased coverage. Discovery is not a hit promise. A stale selected range falls back to the valid engine-owned origin. This compiles declared semantic requirements into deterministic page demand. It does not analyze arbitrary model graphs, prove the model’s mathematics, predict future tokens, authorize reclaim or enable automatic hybrid warming. General retention and physical planning remain future work; this increment has no measured latency claim. See the known-range example and ownership lessons.

Physical bytes may be shared across vLLM and SGLang only when their StateFormat values are compatible. Sharing the core and policy never implies blind cross-framework byte reuse.

Framework adapters

The adapters resolve a shared versioned identity at startup and translate native hashes and GPU layouts into the cache API. The manager binds registered storage geometry and uses StateKey across tiers. Supported layouts use the common recovery contract:

ConcernvLLMSGLang
Prefix identityRequest.block_hashesRadix page hashes
Local GPU pagesvLLM block IDs + CUDA IPCRadix page indices + CUDA IPC on the direct path
Host pagesOrbitKV-owned pinned blocksOrbitKV-owned pinned blocks
Hybrid stateFull + SWA + aligned recurrent groups; shared demand and validationFull + SWA + recurrent/conv; shared demand and validation
LifecycleKVConnector callbacksRadix-cache events

Adapters do not decide which component set is a legal recovery point. That logic belongs in the common recovery contract.

orbitkv-core

The current core provides content-addressed sealed blocks, NUMA-aware pinned memory, leases, LRU/TinyLFU admission, SSD, remote fetch, and session cleanup. DRAM, SSD and the directory share orbitkv-state::StateKey; registration binds model identity to stored layout before Query/Publish. Engine page IDs remain raw, and generation-qualified page types and complete recovery proofs are not yet enforced. See state identity for fingerprint configuration.

Transfer and backing domains

The native physical domains are:

  • framework GPU pages;
  • shared or OrbitKV-owned pinned DRAM;
  • local SSD;
  • remote OrbitKV replicas over Mooncake-selected RDMA or TCP.

Mooncake Transfer Engine is the sole remote-movement backend in this codebase. It contributes Segment/BatchTransfer, multi-NIC topology selection, endpoint pooling, and rail failover. Mooncake Store Master is not OrbitKV’s semantic authority: bundle completeness, leases, generations, and planning remain in OrbitKV.

SGLang integration

Direct GPU linker and compiled recovery

orbitkv.sglang.linker.OrbitKVLinker is registered through SGLang’s plugin entry point and selected by --radix-cache-backend orbitkv together with --enable-unified-cache-external-linker. The latter is required for SGLang’s scheduler to submit GPU restores and drain linker completions. It uses UnifiedCacheLinker callbacks to look up radix page hashes, pin SGLang-owned GPU slots during asynchronous saves and loads, and transfer bytes through the same Cache Manager API as vLLM. Each scheduler rank registers its local GPU KV buffers through CUDA IPC. A model-, rank-, and layout-scoped namespace prevents incompatible byte reuse. Full attention, Full + SWA, Full + recurrent/conv and their combined layout have explicit recovery rules. Convolution and recurrent tensors share one sealed checkpoint group; SWA has independent page coverage. SGLang retains authority over HBM allocation, request-state copy-on-write and prefix-tree nodes. Hybrid lookup discovers group positions without reading payloads and preserves all legal boundaries until rank intersection. Rust reads only the selected compiled ranges; SGLang admits the hit after every rank holds complete leases. The hybrid recovery contract describes the pinned-release component bridge and unsupported representations.

Both DRAM and SSD recovery are GPU-validated at TP=1. SGLang’s general plugin admission hook retains pending requests in the queue and consumes the ready result on a subsequent match. vLLM reports unresolved lookups through its own connector scheduler contract. The original SSD readiness failure and successful follow-up remain in SSD results. The first bounded queued-warming path is implemented; cost selection remains in state demand and transfer planning.

Future: Radix lifecycle bridge for routing

Publish prefix materialization, match, release, promotion, demotion, and removal events from RadixAttention. OrbitKV uses the events to maintain a global replica index and estimate next touch. It does not maintain a competing radix tree.

Future: generation-safe page references

Adapters pass generation-qualified references for engine-owned HBM pages. OrbitKV validates the registration session and page generation before copying, while the engine still allocates and reuses its HBM slots. OrbitKV can assign handles to its own external replicas without taking over the GPU allocator.

Safety invariant

A physical generation may be reused only when both conditions hold:

SemanticDead(page, semantic_frontier)
and
ExecutionComplete(page, execution_frontier)

Semantic death proves that no future legal execution can read the state. Execution completion proves that no submitted CUDA, SSD, or network operation still references the generation. A lease or refcount supplies execution evidence; it does not by itself prove semantic death.

The descriptor arena validates its slot generation and Cache Manager session epoch, and vLLM pins save-source blocks until Publish returns. These checks do not yet validate a framework HBM page’s reuse generation. Publish and Restore still carry block IDs; unused page-reference types are not exposed as guarantees. Generation enforcement requires page-lifecycle information from the adapter and destination ownership through terminal DMA completion.

Multi-node cache path and deployment

orbitkv-catalog is an embedded library served on each distributed Manager’s peer endpoint. All Managers agree on an immutable catalog host set in etcd. Sixteen fixed logical shards are assigned by equal-weight rendezvous hashing; member loss does not change placement. Cached member snapshots resolve each assigned Node ID to a current endpoint and runtime UUID. Ordinary block operations perform no etcd I/O.

Managers asynchronously synchronize independently ordered DRAM inventory streams per shard. Bounded snapshots and deltas reconstruct lost evidence; incomplete replacement views stay hidden until commit. After a local miss, the requester checks its bounded positive candidate index and queries only missing shards. It plans source spans and obtains exact runtime/residency authorization before Mooncake reads bytes into pinned DRAM, then restores them through the same engine API. The destination also advertises its newly resident replicas.

Catalog and source control use gRPC; Mooncake carries KV bytes. Mooncake’s P2P handshake provides transport metadata rather than KV ownership. Each catalog shard currently has one metadata copy, so losing a host makes those cold lookups unavailable until it returns and inventories replay. Other shards and valid cached candidates remain usable. etcd membership gates new remote admission; local DRAM/SSD operations continue through coordinator loss.

Source transfer timeout reclamation still lacks transport revocation qualification. Caller cancellation retains buffers and source holds through blocking completion, but this does not prove safe source failure or partitions. See the implemented protocol and limits.

The next stages add replicated placement generations, controlled handoff, subscriptions and remote SSD. These are target features in the diagram below. The distributed cache design defines the acceptance gates. A later KV-aware router can consume replica summaries and engine load events without entering the transfer path. Metadata replicas do not imply KV payload replicas or general object-store CAS semantics.

host A                                           host B
engine HBM                                      engine HBM
    | UDS + iceoryx2 / CUDA IPC                      | UDS + iceoryx2 / CUDA IPC
Cache Manager A  <---- Mooncake KV bytes ---->  Cache Manager B
  DRAM / SSD · local candidate index              DRAM / SSD · local candidate index
  catalog shards  <---- replicated metadata ---> catalog shards
         \________ etcd membership/placement _________/

catalog summaries ----> future KV-aware router <---- engine load/events

Planning direction

The proposed state demand and transfer planner describes engine readiness signals, recovery boundaries, measured local/peer paths and prefetch timing. Shared Rust cost observations and resource accounting support different deployment contracts.

The first structural refactor adds Core planning/: bounded replica records separate medium from acquisition evidence, SSD planning revalidates exact versions before pinning, and peer plans own source segmentation and rejected-evidence updates. Default execution remains unchanged. The fuller owner/resource endpoint, route and resource-reservation contract remains planned; current discovery still returns positions to the engine. A common endpoint does not give the Manager allocation or eviction authority over engine HBM. GPUDirect RDMA remains a TE capability requiring valid GPU endpoints, not a new cache tier. See the target Core layout and route cost contract. Local path selection comes first; distributed qualification proceeds alongside it. These are design proposals, not capabilities implied by current cache hits.

The planner compares legal recovery boundaries and the exact required_ranges for each, then chooses available sources and executable paths. The longest prefix or smallest stored representation need not minimize request latency. For whole-restore admission, estimate from one decision point:

state_ready = resource wait + restore critical path + completion visibility
first_token = max(engine_admission, state_ready) + remaining_prefill

Observe queue/service time, encoded and physical bytes, fragmentation, device contention and prediction error. Enforce capacity, quality and latency limits; overlapping durations cannot simply be added. Retention/write admission and transfer scheduling share observations but remain separate decisions.

Mooncake TE remains the remote byte engine. Candidate discovery, source authorization and both peers’ budgets stay with OrbitKV. DP can propose bounded fallback to recomputation; current-request P/D handoff needs explicit recovery on failure. TP/PP add rank/stage completion dependencies. These dimensions can compose within a deployment; roles belong to engine instances and operations, not a single global Manager mode. HBM allocation stays with each engine.

Reuse Dynamo’s worker selector for request placement. Dynamo v1.5.0 provides an independent Rust router crate and a selection service; its runtime is an optional dependency of the crate. The selected Cache Manager then revalidates replicas and constructs the leased physical plan: source, restore or recompute proposal, staging budget, transfer deadline, and completion dependencies. The engine owns execution admission and HBM allocation. A routing load reservation does not replace a transfer lease.

The Dynamo integration boundary and implementation stages specify the reusable components, pending engine-interface dependency, and acceptance gates. No router dependency is introduced into the current single-node core.

Ownership boundary by milestone

MilestoneFramework ownsOrbitKV owns
M0local page identity and executionexternal replicas and current data plane
M1GPU pagesPinned DRAM/SSD replicas and direct GPU restore
M2GPU pagesCompiled page demand, common recovery validation and transfer operations
M2.5GPU pages and executionrecoverable replica catalog and remote cache fetch
M3execution and local page identityKV-aware routing and restore plans
M4HBM allocation, physical GPU page IDs, and executionexternal replica handles, validated GPU references, and transfer fences
M5+executionsemantic lifetime and compiled physical plans