In a world increasingly shaped by artificial intelligence, a new study has delivered what initially sounds like a profound blow to human exceptionalism.
AI, it seems, “understands” emotion better than we do.
Published on May 21st in the journal Communications Psychology, research by scientists from the University of Geneva (UNIGE) and the University of Bern (UniBE) suggests that common large language models (LLMs) are remarkably adept at navigating emotionally charged situations, at least on paper.
The researchers put a range of widely-used LLMs, including powerhouses like ChatGPT-4, Gemini 1.5 Flash, and Claude 3.5 Haiku, through a battery of established emotional intelligence (EI) tests.
These weren’t casual quizzes but validated psychological instruments designed to gauge a subject’s ability to perceive, use, understand, and manage emotions.
The results were startling: the AI models selected the “correct” response – based on the consensus of human experts – in 81% of cases, significantly outperforming the average human score of 56%.
As if to underscore their prowess, when ChatGPT was tasked with generating entirely new test questions, human assessors found these AI-created queries to be just as challenging and valid as the originals, demonstrating a “strong” correlation.
The headline conclusion was undeniable: AI appears to grasp emotions with a proficiency that surpasses its creators.
Yet, as with many grand pronouncements about AI, the deeper story is often far more nuanced.
When Live Science sought the perspectives of various experts, a consistent theme emerged: the methodology, while academically sound, bears little resemblance to the messy, unpredictable reality of human emotional interaction.
The EI tests used were, crucially, multiple-choice.
This structured, quantitative environment is precisely where AI, with its unparalleled capacity for pattern recognition and data processing, excels.
“It’s worth noting that humans don’t always agree on what someone else is feeling, and even psychologists can interpret emotional signals differently,” observed Taimur Ijlal, an expert in finance and information security.
He cautioned against conflating statistical success with genuine insight.
“So ‘beating’ a human on a test like this doesn’t necessarily mean the AI has deeper insight. It means it gave the statistically expected answer more often.”
This distinction is critical.
AI systems are masters at identifying patterns, especially when emotional cues conform to a recognizable structure, be it facial expressions in an image or linguistic signals in text.
But, as Nauman Jaffar, Founder and CEO of CliniScripts, an AI-powered documentation tool for mental health professionals, pointed out, “equating that to a deeper ‘understanding’ of human emotion risks overstating what AI is actually doing.”
The core of the expert skepticism lies in the difference between pattern recognition and true understanding.
Emotions are not static, multiple-choice options; they are fluid, context-dependent, and often contradictory.
Jason Hennessey, founder and CEO of Hennessy Digital, drew a parallel to the “Reading the Mind in the Eyes Test,” where AI has shown promise.
However, Hennessey noted that even minor variables like changes in lighting or cultural context can cause “AI accuracy [to drop] off a cliff.” This highlights a fundamental limitation: AI performs well on tests about emotional situations, not in the heat of the moment, where humans experience them with all their inherent complexities and ambiguities.
“Does it show LLMs are useful for categorizing common emotional reactions?” mused Wyatt Mayham, founder of Northwest IT Consulting.
“Sure. But it’s like saying someone’s a great therapist because they scored well on an emotionally themed BuzzFeed quiz.”
Despite this healthy skepticism, there’s a compelling counter-caveat that brings the discussion back to the practical realm.
While AI may not possess the same intuitive, lived understanding of emotion as humans, its ability to process vast amounts of data and identify subtle cues can yield tangible benefits, even in highly charged real-world scenarios.
Consider Aílton, a conversational AI assistant deployed to over 6,000 long-haul truck drivers in Brazil.
This multimodal WhatsApp assistant, developed by Marcos Alves, CEO & Chief Scientist at HAL-AI, uses voice, text, and images to interact with drivers in real time.
Alves claims Aílton identifies stress, anger, or sadness with approximately 80% accuracy – a remarkable 20 points higher than its human counterparts.
One striking example illustrates Aílton’s practical efficacy.
After a colleague’s fatal crash, a distraught driver sent a 15-second voice note.
Aílton responded instantly and appropriately, offering nuanced condolences, providing mental health resources, and automatically alerting fleet managers.
This was not a multiple-choice question in a sterile lab but a raw, unfiltered expression of grief on the open road.
Alves readily acknowledges the limitations of lab studies, stating, “Yes, multiple-choice text vignettes simplify emotion recognition. Real empathy is continuous and multimodal.But isolating the cognitive layer is useful. It reveals whether an LLM can spot emotional cues before adding situational noise.”
He posits that LLMs’ capacity to absorb billions of sentences and thousands of hours of conversational audio allows them to encode “micro-intonation cues humans often miss.” The journey from academic triumph in a controlled environment to practical application in the unpredictable world is fraught with challenges.
While the recent study offers a fascinating glimpse into AI’s growing ability to process and respond to emotional signals, the consensus remains that “understanding” in the human sense is still largely beyond its grasp.
Yet, the case of Aílton suggests that even if AI’s emotional intelligence is a form of highly sophisticated pattern matching rather than genuine empathy, its ability to offer “scalable empathy at scale” could prove invaluable in a world increasingly in need of support, connection, and nuanced response, even if delivered by a machine.
The question then shifts from whether AI truly “feels” to what practical good its remarkable mimicry of feeling can achieve.
-
Frank DiBernardo handles LNGFRM's Foodie and Miscellaneous writing tasks. He's always getting ideas from users, so don't be afraid to send an email to the editor.