Providers Have Brought Kimi K3 Model to U.S. Servers

Three inference platforms now offer cost-effective, domestic access to Moonshot AI's massive 2.8 trillion-parameter model.

Updated on Sept. 21, 2026 in Artificial Intelligence

Isometric editorial illustration showing a high-performance computer accelerator module inside a modular server rack, representing data center technology.
Modal, Fireworks AI, and Baseten have launched U.S.-based hosted inference services for Moonshot AI’s massive Kimi K3 model, providing enterprises with cost-effective, domestic access. AI Illustration. Upload story photo >

Live Poll

Would you trust a third-party provider to host your company's sensitive data for AI processing?

Modal, Fireworks AI, and Baseten have launched U.S.-based hosted inference services for Moonshot AI's Kimi K3 model. This release provides enterprises with a compliant, domestic endpoint for the 2.8 trillion-parameter system.

Why it matters

The availability allows U.S. companies to leverage large-scale models while meeting strict data sovereignty requirements and bypassing the hardware constraints faced by developers in China. This shift democratizes access to high-performance inference through OpenAI-compatible APIs at roughly one-tenth the cost of direct model access.

Kimi K3 utilizes a Mixture-of-Experts architecture with 896 experts, activating 104 billion parameters per forward pass. The model requires 1.4 terabytes of storage at MXFP4 quantization and supports a 1 million token context window.

The players

Moonshot AI

A Beijing-based developer of large language models known for high-context architectures.

Modal

A serverless cloud platform specializing in high-performance GPU-accelerated AI inference.

Fireworks AI

A platform provider focused on optimizing the deployment and serving of open-weight large language models.

Baseten

An infrastructure firm that provides scalable tools for deploying and monitoring production machine learning models.

The details

The model architecture relies on a Mixture-of-Experts system—a method that activates only a fraction of the total parameters for each query to save compute resources. Modal, Fireworks AI, and Baseten host the model on Nvidia GB300 and AMD MI350X/MI355X accelerators. To maintain high throughput, Modal employs DFlash, a custom speculative decoder—a software technique that predicts multiple future tokens simultaneously—to hit speeds of 460 tokens per second.

Timeline

  1. July 27, 2026: Moonshot AI officially released the Kimi K3 model.

The Tech Race

This deployment marks a departure from traditional reliance on local hardware, instead utilizing the U.S. cloud ecosystem to bypass regional GPU shortages. The availability of Kimi K3 via domestic providers directly challenges the market dominance of models restricted by U.S. export controls.

Enterprises can now integrate Kimi K3 into existing workflows using OpenAI-compatible APIs without needing to manage the 1.4 terabyte weight files manually. Pricing is standardized at $3 per million input tokens, $0.30 per million cached tokens, and $15 per million output tokens for all providers.

The takeaway

The move demonstrates that massive models can achieve wider utility through specialized hosting platforms that manage the heavy computational requirements. Watch for future benchmarks as developers stress-test the 1 million token context window in production environments.

Further reading

For more information on the evolving landscape of model serving, visit Artificial Intelligence.

Live Poll

Would you trust a third-party provider to host your company's sensitive data for AI processing?