测试视觉语言模型在多元文化中的心理理论能力,发现其仍有显著局限。
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- 构建跨文化心理理论基准数据集,覆盖6类任务与4级复杂度。
- 顶尖模型准确率超93%,但对错误信念推理仍仅19%-83%。
- 模型存在社会偏好偏差,适合评估跨文化社交推理的研究者使用。
心理理论(ToM)——即理解他人信念与意图的能力——是社会智能的核心,但现有视觉语言模型(VLM)评估仍以西方文化为中心。本文提出CulturalToM-VQA,一个包含5,095个视觉情境化心理理论探针的基准数据集,涵盖多样文化背景、仪式与社会规范。该数据集通过前沿私有多模态大模型与人工验证流程构建,覆盖六类ToM任务与四个复杂度层级。我们评估了10个2023-2025年的VLM,发现早期模型表现不佳,而前沿模型准确率显著提升(>93%)。然而,模型在错误信念推理上仍表现薄弱(19%-83%准确率),且区域间性能差距达20%-30%。关键发现:顶级模型存在社会宜人性偏差——系统性偏向语义积极的答案选项。消融实验表明,部分前沿模型严重依赖参数化社会先验,常默认采用安全对齐预测。尽管思维链提示对旧模型有效,对新模型增益甚微。本工作为跨文化社交推理提供测试平台,强调即便架构进步,实现鲁棒、视觉锚定的理解仍是开放挑战。
原文摘要 · Abstract (English)
Theory of Mind (ToM) - the ability to attribute beliefs and intents to others - is fundamental for social intelligence, yet Vision-Language Model (VLM) evaluations remain largely Western-centric. In this work, we introduce CulturalToM-VQA, a benchmark of 5,095 visually situated ToM probes across diverse cultural contexts, rituals, and social norms. Constructed through a frontier proprietary MLLM, human-verified pipeline, the dataset spans a taxonomy of six ToM tasks and four complexity levels. We benchmark 10 VLMs (2023-2025) and observe a significant performance leap: while earlier models struggle, frontier models achieve high accuracy (>93%). However, significant limitations persist: models struggle with false belief reasoning (19-83% accuracy) and show high regional variance (20-30% gaps). Crucially, we find that SOTA models exhibit social desirability bias - systematically favoring semantically positive answer choices over negative ones. Ablation experiments reveal that some frontier models rely heavily on parametric social priors, frequently defaulting to safety-aligned predictions. Furthermore, while Chain-of-Thought prompting aids older models, it yields minimal gains for newer ones. Overall, our work provides a testbed for cross-cultural social reasoning, underscoring that despite architectural gains, achieving robust, visually grounded understanding remains an open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。