Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters. Before public release, it processed nearly a trillion tokens per day internally. That traffic came from RL rollouts, synthetic data generation, evaluations and long-running coding agents.
What is Prime Inference?
Prime Inference is the serving layer of Prime Intellect’s open training stack. The company already ships post-training tools such as prime-rl, verifiers and sandboxes. Serving closes that loop: deployed models generate production traces that can feed back into training. Prime reports its GLM-5.3 endpoint ranks among the fastest on OpenRouter. It also cites a near-zero tool-call error rate and 100% uptime since launch.
- Two modes: serverless endpoints for variable demand, reserved capacity for sustained workloads.
- OpenAI compatible: point any OpenAI SDK at
https://api.pinference.ai/api/v1(docs). - Uptime: automatic failover across datacenters routes traffic to healthy deployments.
- Hardware: NVIDIA Blackwell today, with Vera Rubin listed as coming soon.
- Billing: unified billing with team-level usage tracking. Per-model pricing is not yet fully published in the docs.
How the serving stack works
The stack combines NVIDIA Dynamo, vLLM, Mooncake and FlashInfer. It was built with Inferact and NVIDIA, and fixes are contributed upstream.
The target workload is agentic. A typical agent turn adds about 6K tokens to a 140K-token prompt. Prime benchmarks this mix with SemiAnalysis AgentX, and injected cold arrivals.
Prefill/decode disaggregation: Prefill and decode run on separate GPU groups. Dynamo handles routing, and vLLM runs the model on each group. Decoders pull computed KV through NIXL. Prime reports nearly 40% lower p90 inter-token latency in its tests.
Cache-aware routing:Dynamo’s KV-aware router weighs cached prefix overlap against queued work. Sessions stay on the same decoder between turns. Mooncake adds a second KV tier in host DRAM.
GLM-5.3 on GB200 NVL72: the numbers
The interactivity target was 100 end-to-end tokens per second per user. At that bar, a 1:4 prefill/decode ratio served the most users. It reached 66 sessions per prefill group at 101 tok/s per user and 100 output tok/s per GPU.
- DEP8 prefill topology: roughly 5x more usable prefix-cache capacity than TEP8 on the same hardware.
- Smaller prefill budget: halving tokens per step from 8K to 4K per GPU cut median queue wait from 550 ms to 110 ms. Median time to first token fell about 20%.
- NVFP4 KV compression: each MLA cache row shrank from 576 to 352 bytes. Cached tokens per decoder rose from 1.09M to 1.63M.
- Native sparse-MLA kernel: about 12.0 μs at 15 query tokens, versus 17.7 μs staged and 13.7 μs FP8. Prime notes this is workload specific.
- BLHNC KV layout: transfer descriptors fell from 19,559 to about 1,940. Mean transfer time dropped from 146 ms to 78 ms.
Reliable tool calls
Agents fail when tool calls carry wrong names or broken arguments. Prime Intellect’s team contributed a structural-tag builder to Dynamo for GLM’s tool format. vLLM then uses xgrammar to mask tokens that violate the tool schema. The team also fixed parsing bugs, including < being decoded into < inside code.
Interactive explainer
Prime Inference vs closest competitors
| Feature | Prime Inference | Together AI | Fireworks AI | Baseten |
|---|---|---|---|---|
| Serverless GLM-5.3 | Yes (source) | Yes (source) | Yes (source) | Yes (source) |
| GLM-5.3 price, input / output per 1M tokens | Not yet published in docs | $1.40 / $4.40 (source) | $1.40 / $4.40 (tracker) | $1.40 / $4.40 (tracker) |
| Dedicated or reserved capacity | Reserved capacity; 1-click dedicated deploys on roadmap | Dedicated Model endpoints (source) | On-demand dedicated GPU deployments (source) | Dedicated GPU deployments (source) |
| OpenAI-compatible API | Yes | Yes | Yes | Yes |
| Batch inference | On roadmap | Yes (source) | Not compared here | Not compared here |
| Disclosed serving stack | Open source: Dynamo, vLLM, Mooncake, FlashInfer | Together inference research stack | Fireworks serving stack | Baseten Inference Stack (source) |
Competitor prices verified October 2, 2026. Tracker figures come from ComputePrices, a third-party price tracker.
Key Takeaways
- Prime Inference is live with serverless and reserved serving for open models.
- GLM-5.3 runs on GB200 NVL72 with Dynamo, vLLM, Mooncake and FlashInfer.
- 1:4 prefill/decode served 66 sessions per prefill group at 101 tok/s per user.
- NVFP4 KV cache lifted capacity from 1.09M to 1.63M tokens per decoder.
- Batch inference and 1-click dedicated deploys are next on the roadmap.
Check out the technical details, docs and the announcement on X. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models appeared first on MarkTechPost.
