Even GPT-5 Failed This Human Attention Test
New scientific research reveals that even advanced models like GPT-5 struggle with the Stroop test as task length increases.
Velocity
How fast coverage is spreading — measured hourly from article rate × source diversity. How this works →
The brief
A recent scientific study has identified a significant limitation in the cognitive capabilities of advanced artificial intelligence. According to reports from SciTechDaily and Caliber.Az, these AI models, including the high-profile GPT-5, have failed a specific human attention test known as the Stroop test. The findings indicate that the performance of these models degrades specifically as the length of the task increases, suggesting a failure to maintain a consistent level of attention or processing accuracy over extended sequences of information. Coverage from Baku.ws emphasizes that scientists have now identified what they describe as the main weakness of artificial intelligence through this specific testing methodology.
SciTechDaily further highlights that the failure occurs even in the most advanced iterations of large language models, such as GPT-5, which were previously expected to handle such cognitive tasks with greater ease. Caliber.Az provides the specific detail that the failure is directly linked to the increasing length of the task, meaning the models struggle more as the sequence of the Stroop test grows longer. To understand why this matters, it is necessary to recognize that the Stroop test is a standard psychological measure used to evaluate human attention and processing speed. It typically requires a subject to ignore a distracting stimulus to focus on a specific task, such as naming the color of a word when the word itself spells a different color.
The fact that AI models fail this test as the task expands indicates a fundamental gap between synthetic processing and human-like attention mechanisms, which has broader implications for how AI handles complex, long-form cognitive instructions. Moving forward, observers will be looking for how researchers address this identified weakness in AI architecture. Based on the reports from Caliber.Az, Baku.ws, and SciTechDaily, the primary focus remains on the correlation between task length and model failure. Future developments will likely center on whether subsequent updates to GPT-5 or other models can overcome this specific attention barrier or if this limitation remains an inherent characteristic of current artificial intelligence design.
Synthesized by PULSE from the headlines below under a strict no-invention contract. ✓ fact-checked: all claims supported by sources Updated 48d ago.
Quick answers
Which AI model specifically failed the attention test?
According to SciTechDaily, GPT-5 failed the human attention test.
What is the specific test the AI models struggled with?
The models were tested using the Stroop test, which measures human attention.
Under what conditions do the AI models tend to fail the test?
Caliber.Az reports that AI models fail the Stroop test as the length of the task increases.
Coverage (3)
- Study finds AI models fail Stroop test as task length increases Caliber.Az · 92d ago
- Scientists have identified the main weakness of artificial intelligence Baku.ws · 92d ago
- Even GPT-5 Failed This Human Attention Test SciTechDaily · 92d ago
Topics
Related trends
Oracle begins a new round of layoffs
Oracle is expanding its ongoing workforce reduction initiative with hundreds of millions in fresh restructuring costs tied to AI transformation.
Microsoft sets limits for future AI models as industry throttles frontier development
Microsoft sets limits for future AI models as the wider industry throttles frontier development.
Microsoft CEO says superintelligence must remain 'under human control'
Microsoft issues its first code of conduct for AI models, emphasizing that superintelligence must stay under human control.
Google Just Dropped a Bombshell on AI Spending
Google releases significant updates regarding artificial intelligence spending and custom chip returns, shifting market focus.
They created Siri before Apple bought it. They’re still waiting for their original vision to come true
Original creators of Siri await their initial vision as new reports highlight payment models and device shifts in artificial intelligence.
Oracle’s $664 billion backlog sends a signal to Dell, HPE
Oracle's massive financial backlog and surging AI demand propel Dell and HPE to the top of the S&P 500.