Open-Weight AI Models Dominated Token Traffic in August

Open-weight models grew to represent 56% of traffic as users shifted to lower-cost alternatives throughout August 2026.

Updated on Oct. 2, 2026 in Artificial Intelligence

Isometric editorial illustration of circular silicon processor wafers in a muted palette, representing the infrastructure of AI compute.
Open-weight AI models accounted for 56% of total token volume on the Vercel AI Gateway in August, as production users prioritized cost-efficient workloads. AI Illustration. Upload story photo >

Live Poll

Do you trust that newer, open-source AI tools are as reliable as expensive proprietary models?

In August 2026, open-weight AI models accounted for 56% of total token volume processed on the Vercel AI Gateway. This represented a rapid shift from December 2025, when these models made up only 7% of traffic.

Why it matters

Production users are increasingly optimizing for cost efficiency, shifting workloads toward lower-cost models to reduce total expenditure. This trend contributed to a 23.2% decline in average token prices across the gateway during the month.

Open-weight models captured 56% of volume but only 14% of total spending, signaling a significant price disparity compared to proprietary alternatives. Meanwhile, high-volume teams processing over 10 million tokens saw a 7.6% decrease in price per token compared to July.

The players

Vercel

A cloud platform provider that operates the Vercel AI Gateway to manage and monitor AI model traffic and costs.

Anthropic

An AI research and development company that captured 64% of total gateway spending in August 2026.

OpenAI

An AI research organization known for its GPT series of large language models and recent launch of GPT-6 Astra.

The details

Users achieved these cost reductions by transitioning workloads from higher-priced models like Fable 5 to lower-cost options such as Opus 5, which saw its share of gateway spend rise to 22.5%. The market remains highly dynamic, with 50% of August token traffic originating from models that have been available for less than three months. Performance also appears tied to rapid deployment cycles, as evidenced by OpenAI's GPT-6 Astra, which captured one in three dollars spent on OpenAI models within 48 hours of its launch.

Timeline

  1. December 2025: Open-weight models comprised 7% of token traffic.

  2. April 2026: Open-weight models grew to 13% of token volume.

  3. August 2026: Open-weight models reached 56% of token traffic.

  4. September 4-16, 2026: GPT-6 Astra and GPT-5.6 Sol processed 27% of OpenAI tokens.

The Tech Race

This data confirms a broader industry shift toward open-weight models as primary drivers of production traffic. The rapid adoption of new releases like GPT-6 Astra indicates that users continue to evaluate new, performant models against cost-optimized incumbents.

Developers and companies managing high-volume AI workloads are likely to see decreasing costs as they rotate to newer, lower-priced open-weight models. These changes require active monitoring of model latency and output consistency, as token expenditures can shift rapidly upon the release of updated competing models.

The takeaway

The move toward open-weight models represents a fundamental strategy shift to maintain output while curbing operational expenditures. Watch the upcoming quarterly spending patterns to see if the 14% share of total spend for open-weight models increases as developers further refine their production stacks.

Further reading

For more information on model performance and adoption, explore our Artificial Intelligence section.

Source note: This article includes information reported by IT Brief Australia.

Live Poll

Do you trust that newer, open-source AI tools are as reliable as expensive proprietary models?