让AI看视频也能懂人心,提升多模态模型的共情推理能力。
Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
- 通过视觉干预向量引导模型关注正确语义,减少语言偏见。
- 在EgoToM数据集上多选题准确率显著提升,生成解释更贴近真实心理状态。
- 适合研究多模态理解、可解释AI与人机协作的开发者和学者。
随着大语言模型(LLMs)的发展,对其推断人类心理状态并展现类人理论心智(ToM)能力的兴趣日益增长。然而,现有大多数ToM评估集中于文本输入,而仅依赖视觉信息的场景却很少被关注。这留下了空白,因为现实世界中的人机交互通常需要多模态理解。此外,许多现有方法将模型视为黑箱,很少探究其在多选题问答中的内部注意力行为,且对模型幻觉在可解释性层面的影响也缺乏研究。为此,我们提出VisionToM——一种面向视觉的干预框架,旨在增强任务感知推理。核心思想是计算与正确语义目标对齐的干预向量,从而引导模型在不同层的视觉特征中调整注意力。该引导减少了模型对虚假语言先验的依赖,提升了多模态语言模型(MLLM)输出的可靠性,并显著改善了问答性能。在EgoToM基准测试(一个用于人类视角、真实世界视频的ToM多选题数据集,包含三种问答设置)上的实验表明,该方法大幅增强了MLLM的理论心智能力。此外,在开放生成任务上的结果还显示,VisionToM使MLLM能够生成更准确捕捉代理心理状态的自由形式解释,推动机器与人类协作更加契合。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). Most existing ToM evaluations, however, are centered on text-based inputs, while scenarios relying solely on visual information receive far less attention. This leaves a gap, since real-world human-AI interaction typically requires multimodal understanding. In addition, many current methods regard the model as a black box and rarely probe how its internal attention behaves in multiple-choice question answering (QA). The impact of LLM hallucinations on such tasks is also underexplored from an interpretability perspective. To address these issues, we introduce VisionToM, a vision-oriented intervention framework designed to strengthen task-aware reasoning. The core idea is to compute intervention vectors that align visual representations with the correct semantic targets, thereby steering the model's attention through different layers of visual features. This guidance reduces the model's reliance on spurious linguistic priors, leading to more reliable multimodal language model (MLLM) outputs and better QA performance. Experiments on the EgoToM benchmark-an egocentric, real-world video dataset for ToM with three multiple-choice QA settings-demonstrate that our method substantially improves the ToM abilities of MLLMs. Furthermore, results on an additional open-ended generation task show that VisionToM enables MLLMs to produce free-form explanations that more accurately capture agents' mental states, pushing machine-human collaboration toward greater alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。