Scale AI Unveils 'Humanity’s Sixth Sense' Benchmark: Humans 93.1% vs AI Peak 53.6%
Scale AI and Elorian launched Humanity’s Sixth Sense (HSS) testing everyday visual reasoning across 522 tasks. Humans reached 93.1%, while top AI scored only 53
On October 7, 2026, Scale AI Labs and Elorian officially released 'Humanity’s Sixth Sense (HSS)', a novel benchmark designed to evaluate the intuitive visual reasoning humans draw upon effortlessly every day. Spanning 522 open-ended tasks across static imagery and short video clips, the evaluation established a human benchmark score of 93.1%. In stark contrast, the highest-performing artificial intelligence model, GPT-6-astra operating at maximum reasoning effort, reached only 53.6%, demonstrating a profound cognitive gap between human perception and frontier multimodal systems.

Image source: Scale AI / Elorian
While leading foundation models regularly surpass human baselines on legal bar examinations and complex text-based reasoning, HSS exposes a fundamental blind spot: frontier AI models routinely fail at basic spatial logic, mechanical constraints, and social intentions that an average adult grasps without conscious deliberation.
522 Open-Ended Tasks Across Four Core Cognitive Domains
The researchers define the benchmark's premise around the implicit sensory deductions people make continuously—spatial layouts, causal dynamics, physical clearances, and social hierarchies—characterizing this everyday intuition as humanity's cognitive "sixth sense."
HSS comprises 522 open-ended free-response tasks derived from 288 static images and 234 video clips, totaling 17.6 hours of visual footage. To eliminate multiple-choice guessing effects and superficial statistical heuristics, every question requires full natural-language synthesis. Submissions are graded against rigorous human evaluation rubrics, strictly disallowing partial credit unless all necessary criteria are satisfied.
The evaluation suite tests four fundamental branches of visual cognition:
- Temporal and Causal Dynamics: Predicting continuous trajectory changes and mechanistic sequences over time (e.g., calculating how many fish in an aquarium will exit the lower edge of the camera frame based on their active swimming vectors).
- Physical and Spatial Logic: Evaluating physical affordances, mechanical tension, and volumetric clearance (e.g., determining whether two additional books will fit into visible shelf gaps, or inferring whether taut chains indicate an active plough towing state).
- Social Understanding: Inferring interpersonal intentions, emotional states, relative power hierarchies, and social norms from situational body language, eye contact, and contextual staging.
- Contextual Reasoning: Reconstructing implicit, unrendered environmental circumstances (e.g., retrodicting which entryway a traveler entered through based on subtle positional cues).
Frontier Leaderboard Results: Top Model Reaches 53.6%, Median at 30.9%
Evaluating 25 leading proprietary and open-weight multimodal models revealed a substantial performance gulf relative to the 93.1% human baseline.
GPT-6-astra, configured with maximum reasoning effort, achieved the highest result on the public leaderboard at 53.6%. However, the median model score across all 25 evaluated architectures languished at just 30.9%, meaning the typical frontier system answered fewer than one in three tasks correctly. Frontier systems including Claude Opus 5.5 and Gemini 3.8 Flash similarly struggled, failing sample tasks involving everyday physical common sense across all three attempts.
The sharpest deficiency appeared in the 'Social Understanding' domain. Among the 25 evaluated models, 21 registered their lowest domain performance in social reasoning, averaging a modest 24.4%. This highlights that discerning nuanced human intentions, gaze direction, and situational hierarchies remains exceptionally difficult to capture purely through large-scale textual pretraining.
Error Taxonomy Across 8,573 Failures: 94% Driven by Perception and Implicit Cue Blindness
To diagnose the underlying failure modes, Scale AI and Elorian conducted an error breakdown across 8,573 incorrect model predictions. Notably, 94% of all errors did not stem from breakdowns in symbolic or deductive reasoning logic.
- Visual Misperception (53%): Models misread or distorted unambiguous spatial and visual evidence directly visible in the frame.
- Implicit Inference Failure (41%): Models failed to extract unrendered but common-sense physical and situational implications implied by the scene.
These findings demonstrate that while modern multimodal architectures possess sophisticated linguistic and logical chains of thought, they lack the grounded sensory perception necessary to discern which visual cues govern real-world physical and social realities.
Scale AI and Elorian have made the complete benchmark dataset openly available on Hugging Face (ScaleAI/HSS), accompanied by an open leaderboard and research paper. The authors noted that leaderboard standings remain subject to variance as model providers push checkpoint updates and adjust reasoning effort parameters.
Sources
- Scale AI Labs Official X: Humanity's Sixth Sense Announcement
- Scale Labs Research: Humanity's Sixth Sense Paper
- Scale Labs Leaderboard: HSS Official Leaderboard
- Hugging Face Dataset: ScaleAI/HSS Dataset Repository
- Elorian Official Blog: Introducing Humanity's Sixth Sense