Prime Intellect Launched Open Model Inference Platform
The new service utilizes NVIDIA Blackwell hardware and a disaggregated architecture to host open-source models.
Updated on Oct. 3, 2026 in Artificial Intelligence

Live Poll
Do you believe serverless platforms make building and scaling AI applications significantly easier for your projects?
Prime Intellect has launched Prime Inference, a serving platform for open-source AI models that offers serverless endpoints and reserved capacity. The system is currently operational on NVIDIA GB200 NVL72 hardware.
Why it matters
The platform addresses the operational complexity of deploying frontier-scale open models by providing a unified stack for handling inference at volume. It aims to streamline performance for high-token-count workloads in the open-source ecosystem.
The stack reported a 40% reduction in p90 inter-token latency and supports 1.63 million cached tokens per decoder. It maintains 100% uptime while enabling 100 tokens/s for interactive sessions.
The players
Prime Intellect
An infrastructure developer focused on large-scale AI serving stacks and open-model deployment.
NVIDIA
The primary hardware supplier providing the GB200 NVL72 and upcoming Vera Rubin GPU architectures.
The details
The platform architecture separates prefill and decoding tasks, assigning distinct GPU groups to each to optimize compute utilization. It further employs cache-aware routing that balances the cost of prefix overlap against pending requests, utilizing host DRAM as a secondary key-value tier. The system integrates NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer to support OpenAI-compatible SDKs.
Timeline
October 2, 2026: Competitor pricing data was verified for comparison.
October 3, 2026: The Prime Inference platform was publicly launched.
The Tech Race
Prime Intellect is positioning its stack as a high-performance alternative to general-purpose cloud inference providers. The platform directly competes with established model-hosting services by integrating architectural optimizations for the GB200 NVL72.
Developers using OpenAI-compatible SDKs can now access serverless endpoints for frontier open-source models. The platform will expand to include one-click dedicated deployments and batch inference in future updates.
The takeaway
Prime Intellect is moving toward a highly optimized inference model that leverages disaggregated GPU clusters for efficiency. Watch for the upcoming addition of batch inference and Vera Rubin hardware support to gauge the platform's scaling capabilities.
Further reading
For more on the infrastructure supporting current model deployments, read the latest reports in Artificial Intelligence.
Live Poll
Do you believe serverless platforms make building and scaling AI applications significantly easier for your projects?








