Anthropic and OpenAI Models Tied in New AI Benchmark

Epoch AI’s new research-automation test measures how well frontier models perform on specialized organizational tasks.

Updated on Oct. 9, 2026 in Artificial Intelligence

Anthropic and OpenAI Models Tied in New AI Benchmark

Live Poll

Do you trust current AI models to perform complex professional research without human supervision?

Epoch AI has released its Automation Reports benchmark, which assesses how effectively frontier models can execute internal research workflows. Anthropic Claude Fable 5.1 and OpenAI GPT-6 Astra shared the top performance with identical 65% scores across 11 tasks.

Why it matters

As labs look to accelerate their own internal development, this benchmark provides a standardized way to measure if AI can successfully automate complex research work. It establishes a performance baseline for the latest generation of models in specialized scientific and analytical environments.

The evaluation measures models across 11 distinct research tasks, with leaders Claude Fable 5.1 and GPT-6 Astra scoring 65%. Other results included Grok 4.6 at 59%, Qwen 3.8 Max at 53%, Kimi K3 at 52%, and Gemini 3.8 Flash at 42%, marking a 20-point spread between the top and bottom.

The players

Epoch AI

An AI research organization focused on forecasting the progress and impact of frontier machine learning models.

Anthropic

A developer of AI models focused on safety and reasoning capabilities, known for the Claude series.

OpenAI

A research laboratory that builds large-scale transformer models like the GPT line.

The details

The benchmark relies on human graders who score model output against quality rubrics derived from Epoch AI’s actual daily research operations. By testing models on these specific workflows, the researchers aim to quantify how well automated systems can assist or replace researchers in technical tasks. The results reflect each model's capacity to handle professional-grade analytical duties rather than generic conversational prompts.

Timeline

  1. October 7, 2026: Epoch AI released the InnovationEval benchmark.

  2. October 8, 2026: Epoch AI released the Automation Reports.

The Tech Race

This report follows a broader trend of establishing specialized benchmarks to differentiate between frontier models. It sits alongside Epoch AI’s InnovationEval, further defining the race to automate high-level scientific and technical work.

These results provide insight for enterprise users on which models are currently best suited for complex, multi-step research and analytical workflows. Organizations should monitor these scores as a signal for which platforms can reliably automate high-stakes internal tasks.

The takeaway

The performance gap indicates that while top models have reached a plateau at 65% in this benchmark, there remains a significant divergence in capabilities for research automation. Watch for future iterations of this benchmark to see if any model can break past this current score ceiling.

Further reading

For more on the current state of model performance evaluation, see our coverage of Artificial Intelligence.

Live Poll

Do you trust current AI models to perform complex professional research without human supervision?