Princeton Researchers Built Chess-Playing Language Model
The 4-billion-parameter Queen model reached a 2697 Elo rating, signaling new potential for explainable AI decision-making.
Updated on Oct. 6, 2026 in Artificial Intelligence

Live Poll
Do you trust AI agents more when they provide clear explanations for their decisions?
Researchers at Princeton have developed a 4-billion-parameter language model named Queen that achieved a 2697 Elo rating on the ChessBench scale. This research-stage development, detailed in an arXiv paper, demonstrates how language models can be trained to master complex expert systems.
Why it matters
The Queen model provides a new method for enabling language models to explain their internal decision-making processes by bridging expert knowledge with linguistic reasoning. This architecture aims to move beyond standard predictive text, with projections suggesting its application may extend to robotics and computer-use agents.
The 4-billion-parameter Queen model improved from an initial 1782 Elo to 2697 Elo over seven training rounds. Researchers observed no performance plateau during this process, significantly outpacing the Gemini 3.1 Pro Preview, which measures near 1149 Elo on the same benchmark.
The players
Princeton
A research university recognized for its contributions to computer science, advanced algorithmic research, and AI architecture development.
The details
The Queen architecture pairs a chess-specific encoder—a module that converts board states into numerical data—with an instruction-tuned language decoder via cross-attention. The training process utilizes a natural-language version of the Bellman update, a mathematical method for reinforcement learning, to perform self-distillation. This approach allows the model to refine its decision-making logic using textual explanations alongside its board evaluations.
Timeline
October 2, 2026: Research paper describing the Queen model was posted to arXiv.
The Tech Race
The Queen model's performance on the ChessBench evaluation scale establishes a new high-water mark for language models in complex strategy domains. This result directly challenges existing benchmarks set by industry-leading models, indicating a rapid advancement in specialized reasoning capabilities.
This development remains at the research stage and is not currently available for public use or commercial integration. Users should watch for future updates regarding how this self-distillation method might be applied to general-purpose agents and computer-use tools.
The takeaway
The Princeton research suggests that language models can be successfully trained to handle complex decision-making tasks through iterative self-distillation. Observers should track future publications from the team to see if this architecture maintains its performance gains in non-chess domains.
Further reading
Learn more about the latest developments in large-scale model architectures on our Artificial Intelligence page.
Source note: This article includes information reported by Startup Fortune.
Live Poll
Do you trust AI agents more when they provide clear explanations for their decisions?









