arXiv:2510.27195cs.CVcs.CL2025-10被引 4

测试大模型在多人社交中识破谎言的能力,发现现有模型表现不佳。

Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions

  • 构建多模态互动真实性评估任务,基于狼人杀游戏数据。
  • 即使GPT-4o也难以可靠区分真话与假话,准确率不足60%。
  • 适合研究多模态推理、可信AI及社交智能的学者参考。

随着AI系统日益融入人类生活,赋予其稳健的社会智能成为关键前沿。其中,辨别真伪是人际互动的核心,依赖语言与非语言视觉线索的复杂交互。然而,在动态多人群体对话中自动识破谎言仍具挑战性。近期强大的多模态大语言模型(MLLMs)在视觉与文本理解上表现优异,使其成为该任务的理想候选。但其在此关键领域的能力尚未被量化。为此,我们提出新任务——多模态互动真实性评估(MIVA),并构建一个源自社交推断游戏《狼人杀》的新型多模态数据集。该数据集包含同步视频与文本,并为每条陈述提供可验证的真实标签。我们建立全面基准,评估顶尖MLLMs性能,结果显示显著差距:即便强大模型如GPT-4o也难以可靠区分真假。失败模式分析表明,这些模型未能有效将语言与视觉社交线索对齐,可能因过度保守的对齐策略而影响判断,凸显亟需新方法以构建更具洞察力与可信度的AI系统。

原文摘要 · Abstract (English)

As AI systems become increasingly integrated into human lives, endowing them with robust social intelligence has emerged as a critical frontier. A key aspect of this intelligence is discerning truth from deception, a ubiquitous element of human interaction that is conveyed through a complex interplay of verbal language and non-verbal visual cues. However, automatic deception detection in dynamic, multi-party conversations remains a significant challenge. The recent rise of powerful Multimodal Large Language Models (MLLMs), with their impressive abilities in visual and textual understanding, makes them natural candidates for this task. Consequently, their capabilities in this crucial domain are mostly unquantified. To address this gap, we introduce a new task, Multimodal Interactive Veracity Assessment (MIVA), and present a novel multimodal dataset derived from the social deduction game Werewolf. This dataset provides synchronized video, text, with verifiable ground-truth labels for every statement. We establish a comprehensive benchmark evaluating state-of-the-art MLLMs, revealing a significant performance gap: even powerful models like GPT-4o struggle to distinguish truth from falsehood reliably. Our analysis of failure modes indicates that these models fail to ground language in visual social cues effectively and may be overly conservative in their alignment, highlighting the urgent need for novel approaches to building more perceptive and trustworthy AI systems.

多模态社会智能谎言检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。