arXiv:2511.16221cs.CVcs.CL2025-11被引 5

测试大模型读不懂社交谎言,发现其难以判断真实与虚假。

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

  • 构建多模态欺骗评估任务MIDA,同步视频与文本带真实标签。
  • 12个主流模型表现不佳,最强如GPT-4o也难准确分辨谎言。
  • 提出社交思维链与动态社会认知记忆,提升模型社交推理能力。

尽管具备先进推理能力,当前主流多模态大模型(MLLMs)明显缺乏人类智能的核心能力:在复杂社交互动中‘读取氛围’并识别欺骗行为。为系统量化这一缺陷,我们提出新任务Multimodal Interactive Deception Assessment(MIDA),构建首个同步视频与文本、带可验证真值标签的多模态数据集。我们对12个先进开源与闭源MLLM进行全面基准测试,揭示显著性能差距:即使强大模型如GPT-4o也难以可靠区分真相与谎言。失败分析表明,这些模型未能有效将语言与多模态社交线索关联,且缺乏对他人知识、信念或意图的建模能力,凸显构建更敏锐可信AI的紧迫性。为此,我们设计了社交思维链(SoCoT)推理流程与动态社会认知记忆(DSEM)模块,实验显示该框架在该挑战任务上取得性能提升,为实现真正类人社交推理提供了可行路径。

原文摘要 · Abstract (English)

Despite their advanced reasoning capabilities, state-of-the-art Multimodal Large Language Models (MLLMs) demonstrably lack a core component of human intelligence: the ability to `read the room' and assess deception in complex social interactions. To rigorously quantify this failure, we introduce a new task, Multimodal Interactive Deception Assessment (MIDA), and present a novel multimodal dataset providing synchronized video and text with verifiable ground-truth labels for every statement. We establish a comprehensive benchmark evaluating 12 state-of-the-art open- and closed-source MLLMs, revealing a significant performance gap: even powerful models like GPT-4o struggle to distinguish truth from falsehood reliably. Our analysis of failure modes indicates that these models fail to effectively ground language in multimodal social cues and lack the ability to model what others know, believe, or intend, highlighting the urgent need for novel approaches to building more perceptive and trustworthy AI systems. To take a step forward, we design a Social Chain-of-Thought (SoCoT) reasoning pipeline and a Dynamic Social Epistemic Memory (DSEM) module. Our framework yields performance improvement on this challenging task, demonstrating a promising new path toward building MLLMs capable of genuine human-like social reasoning.

多模态社交推理欺骗检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。