AI Models Lag Humans on New Social Reasoning Benchmark

The Humanity's Sixth Sense benchmark reveals AI models currently struggle to interpret basic social and physical cues.

Updated on Oct. 9, 2026 in Artificial Intelligence

AI Models Lag Humans on New Social Reasoning Benchmark

Live Poll

Do you trust current AI models to perform tasks requiring social intuition?

Scale AI and Elorian have released the Humanity's Sixth Sense benchmark, a new set of 522 visual tasks designed to test machine reasoning. The study shows that leading AI models currently trail significantly behind human performance in interpreting social and physical contexts.

Why it matters

This benchmark quantifies the gap between current AI capabilities and human social intuition, highlighting a persistent weakness in models' ability to infer emotional states and social rules. It serves as a metric for tracking how well AI can navigate the complexities of real-world human interactions.

Tested against a cohort of 25 AI models, the benchmark utilizes 288 images and 234 video clips to probe social and physical reasoning. While OpenAI's GPT-6-astra reached the top score of 53.6%, the average accuracy on social understanding tasks remained at just 24.4%.

The players

Scale AI

A data infrastructure firm specializing in labeling and curating large datasets for machine learning model training.

Elorian

A research entity focused on evaluating and benchmarking the reasoning capabilities of artificial intelligence systems.

OpenAI

A leading AI developer known for its series of large language and multimodal models, including the GPT-6-astra tested here.

The details

The benchmark tasks require models to process human-written prompts alongside 17.6 hours of total visual footage. Models must infer spatial layouts, social rules, and emotional states from the provided data. This process relies on theory of mind—the cognitive ability to attribute mental states, such as beliefs or intentions, to oneself and others—to interpret the visual input accurately.

Timeline

  1. October 7, 2026: Scale AI and Elorian released the Humanity's Sixth Sense benchmark.

The Tech Race

This release follows a broader industry push to move beyond static, language-heavy benchmarks like MMLU toward dynamic visual reasoning tests. By forcing models to interpret social cues from video, the benchmark attempts to measure performance in areas where current large multimodal models have historically faltered.

For users and developers, this data indicates that current AI tools still require significant human oversight for complex social decision-making. These findings will influence how developers prioritize future model training, focusing on physical common sense and emotional reasoning in upcoming iterations.

The takeaway

The performance gap between AI and humans on this benchmark suggests that true social fluency remains a significant hurdle for current architectures. Watch for future model releases to see if developers can bridge this discrepancy in subsequent iterations of visual reasoning tasks.

Further reading

For more on the current state of model evaluation, see our latest coverage in Artificial Intelligence.

Source note: This article includes information reported by Startup Fortune.

Live Poll

Do you trust current AI models to perform tasks requiring social intuition?