Princeton Researchers Built Chess-Playing Language Model

The 4-billion-parameter Queen model reached a 2697 Elo rating, signaling new potential for explainable AI decision-making.

Updated on Oct. 6, 2026 in Artificial Intelligence

Isometric editorial illustration of floating cubic neural nodes and connecting conduits, representing complex artificial intelligence decision pathways.
Princeton University researchers have developed 'Queen,' a 4-billion-parameter AI model that achieved a 2697 Elo rating, advancing explainable decision-making in machine learning. AI Illustration. Upload story photo >

Live Poll

Do you trust AI agents more when they provide clear explanations for their decisions?

Researchers at Princeton have developed a 4-billion-parameter language model named Queen that achieved a 2697 Elo rating on the ChessBench scale. This research-stage development, detailed in an arXiv paper, demonstrates how language models can be trained to master complex expert systems.

Why it matters

The Queen model provides a new method for enabling language models to explain their internal decision-making processes by bridging expert knowledge with linguistic reasoning. This architecture aims to move beyond standard predictive text, with projections suggesting its application may extend to robotics and computer-use agents.

The 4-billion-parameter Queen model improved from an initial 1782 Elo to 2697 Elo over seven training rounds. Researchers observed no performance plateau during this process, significantly outpacing the Gemini 3.1 Pro Preview, which measures near 1149 Elo on the same benchmark.

The players

Princeton

A research university recognized for its contributions to computer science, advanced algorithmic research, and AI architecture development.

The details

The Queen architecture pairs a chess-specific encoder—a module that converts board states into numerical data—with an instruction-tuned language decoder via cross-attention. The training process utilizes a natural-language version of the Bellman update, a mathematical method for reinforcement learning, to perform self-distillation. This approach allows the model to refine its decision-making logic using textual explanations alongside its board evaluations.

Timeline

  1. October 2, 2026: Research paper describing the Queen model was posted to arXiv.

The Tech Race

The Queen model's performance on the ChessBench evaluation scale establishes a new high-water mark for language models in complex strategy domains. This result directly challenges existing benchmarks set by industry-leading models, indicating a rapid advancement in specialized reasoning capabilities.

This development remains at the research stage and is not currently available for public use or commercial integration. Users should watch for future updates regarding how this self-distillation method might be applied to general-purpose agents and computer-use tools.

The takeaway

The Princeton research suggests that language models can be successfully trained to handle complex decision-making tasks through iterative self-distillation. Observers should track future publications from the team to see if this architecture maintains its performance gains in non-chess domains.

Further reading

Learn more about the latest developments in large-scale model architectures on our Artificial Intelligence page.

Source note: This article includes information reported by Startup Fortune.

Live Poll

Do you trust AI agents more when they provide clear explanations for their decisions?