Prime Inference: Fast, Reliable Serving for Frontier Open Models Prime's mission is to build frontier open models and the open superintelligence stack for continuously improving agents. We already provide end-to-end post-training infrastructure, from prime-rl and verifiers to sandboxes and RL environments. But the continual learning loop is not complete until a trained model can serve real users, generate new experience, and feed those production traces back into training. We're excited to release Prime Inference today. It covers both serverless endpoints and reserved capacity and offers resilient serving of frontier open-source models on our GPU infrastructure across multiple datacenters. Prime Inference began as the serving platform we needed ourselves. Long before public release, it powered large-scale RL rollouts, synthetic data generation, evaluations, and long-running coding agents, processing nearly a trillion tokens every day just internally. Beyond our own workloads, we've also been serving large-scale customer deployments in production since January. This scale pushed us to optimize for sustained performance, quality, and reliability, rather than benchmark speed alone. Our first public deployment, GLM-5.3, went live on OpenRouter on September 22. It currently ranks among the fastest GLM-5.3 endpoints on OpenRouter, with a near-zero tool-call error rate and 100% uptime since launch. Prime Inference at a glance Low-latency: Our GLM-5.3 endpoint on OpenRouter is continuously evaluated for quality, with production SLAs, security, and privacy built in from the start. Premium infrastructure across data centers: Prime-hosted models run on NVIDIA Blackwell today, with Vera Rubin coming soon. Uptime: Automatic failover across data centers keeps traffic moving to healthy deployments. OpenAI compatible: Connect your existing tools and SDKs using a Prime endpoint and API key. Scaling: Serverless endpoints for variable demand, with reserved capacity for sustained workloads. Cost: Unified billing and team-level usage tracking across models, making inference spend easier to manage. Robust open-source infrastructure: Our stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed in close partnership with Inferact and NVIDIA, with improvements contributed upstream. Get started CLI cURL # uv tool install prime && prime login prime inference chat 'z-ai/glm-5.3' "Write a haiku about KV caches." Or point any OpenAI SDK at https://api.pinference.ai/api/v1 . See the docs for the full API reference. Built for production SLAs Prime Inference separates the public API from the model fleet, so capacity can move, fail, or scale without changing the client endpoint. Our shared circuit breakers let every gateway replica react to failures consistently, while lease-based admission control prevents overload and automatically recovers capacity when a process disappears. Below the software layer, every cluster is continuously monitored, with health checks that reach all the way down to NVLink and InfiniBand. Alerts go to an on-call team staffed 24/7, so GPU failures are caught and repaired quickly instead of slowly degrading service. And because Prime maintains significant overflow capacity, we can reroute traffic and bring up new deployments whenever more capacity is needed. Together, these safeguards keep the service available. The next sections look inside a production deployment through our work serving GLM-5.3 on GB200 NVL72. How we serve production agent traffic Workload A typical agent turn adds about 6K tokens to a 140K-token prompt, reusing most of the conversation history. Under load, these returning sessions run alongside new requests with long, uncached prompts. We benchmark this mix with AgentX from SemiAnalysis, which replays multi-turn agent sessions. Our benchmark harness also injects cold arrivals with long prompts. We measure end-to-end tokens per second per user for interactivity and output tokens per second per GPU for efficiency. Prefill/decode disaggregation On shared GPUs, processing a long prompt can interrupt token generation for existing sessions. Chunked prefill limits these interruptions, but both workloads still compete for GPU time. We run prefill and decode on separate GPU groups. NVIDIA Dynamo handles routing and orchestration, while vLLM runs the model on each group. Once prefill finishes, the decoder pulls the computed KV through NIXL and adds the request to its batch. Chunked prefill on shared GPUs versus separate prefill and decode workers. With Dynamo coordinating separate prefill and decode pools, we reduced p90 inter-token latency by nearly 40% in our tests. Caching and routing Dynamo's KV-aware router chooses a prefill worker based on how much of the prompt it already has cached and how much work is queued there. Workers publish cache updates so the router can track where prefixes are available. We also keep sessions on the same decoder between turns to support KV reuse. Mooncake provides a second cache tier in host DRAM. Prefixes offloaded from GPU memory can be retrieved instead of recomputed, allowing us to retain more conversation history. The router balances cached prefix overlap against queued work. Mooncake holds cached KV outside GPU memory so workers can retrieve it when needed. With this architecture in place, we tuned GLM-5.3 on GB200 NVL72 for three goals at once: interactive speed, model quality, and concurrency. Performance: GLM-5.3 on GB200 NVL72 Long-context agentic serving is as much a cache-management problem as a compute problem. Performance therefore depends on retaining that history, scheduling new work promptly, and moving cached state without interrupting ongoing generation. Our interactivity target was 100 end-to-end tokens per second per user. We tune for the number of concurrent sessions we can support at that speed. We optimized these paths separately: prefill topology and scheduling to reduce time to first token; compressed KV and a fused attention kernel to support low-latency decoding; and a transfer-friendly cache layout to reduce the overhead of moving KV between workers. At the 100 tok/s/user bar, a 1:4 P/D ratio serves the most: 66 sessions per prefill group at 101 tok/s/user and 100 output tok/s per GPU. Technical deep dive For readers who want the engineering details, the rest of this post walks through each optimization in depth, followed by our work on reliable tool calls. The sections below cover our work on topology, scheduling, compressed attention, and KV transfer: Choosing the right topology for prefill and decode Reducing the scheduler bubble on prefill NVFP4 KV compression on FlashInfer Faster NIXL transfers on NVLink with the BLHNC layout Prefill: time to first token Time to first token depends on more than processing the prompt. A request may need to retrieve cached history, wait for admission, compute new tokens, and transfer KV to a decoder. We investigated delays across this path, starting with cache capacity and scheduling. Choosing the right topology We chose DEP8: eight data-parallel attention ranks with expert parallelism across the group. DEP8 distributes requests across eight attention workers while sharing the model's experts across the group. With the MLA cache layout we used, TEP8 replicated each request's KV across all eight ranks, while DEP8 let the ranks cache different requests. Even after accounting for the extra weight memory this requires, we had roughly five times more usable prefix-cache capacity on the same hardware compared to a topology like TEP8, which was also benchmarked. The downside of DEP8 is the DP rank synchronization: each DP rank processes different requests but joins the same all-to-all communication at every MoE layer, so even an idle rank may need to run forward passes to keep up with its peers. Timing breakdown in a DEP8 rank's prefill execution. Its own forward takes 408 ms, followed by roughly 245 ms of dummy work while the other ranks finish. The dispatch overhead would further increase as EP goes wider. We found 1 DEP16 took 17.9% longer than two DEP8 groups, with combine and finalization growing the most. Scheduler bubble and the token budget Cache capacity does not eliminate scheduling delays. We found that a request's cached KV could already be available while the request still waited to enter the running batch. Workers check for completed loads between forward passes. Results then pass through a batch queue and scheduler, where a ready request can miss the current decision and wait another step. Under load, requests already in progress can fill the next prefill batch, delaying admission even when the cached history is ready. This creates a scheduler bubble. To reduce this bubble, we halved the number of prompt tokens processed in each prefill step, from 8K to 4K per GPU. Shorter steps let waiting requests start sooner. On our configuration, median queue wait fell from 550
Prime Inference: Fast, Reliable Serving for Frontier Open Models Prime
Prime Inference: Fast, Reliable Serving for Frontier Open Models Prime's mission is to build frontier open models and the open superintelligence stack for continuously improving agents. We already provide end-to-end post-training infrastructure, from prim
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。