arXiv:2605.22109cs.AIcs.CV2026-05

测试大模型能否真正理解人格而非仅凭表面线索猜得分。

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

论文配图:Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
图 1 · 摘自论文原文
  • 提出新任务GPR,要求模型给出人格评分并附带行为证据链。
  • 构建含1104段视频的MM-OCEAN数据集,支持多维度证据匹配。
  • 发现51%正确评分无真实证据支撑,暴露模型刻板印象问题。

多模态大语言模型在涉及人格判断的人机交互场景中日益普及,但现有评测仅依赖五大性格特质(Big Five)分数预测,无法判断模型是通过行为理解感知人格,还是仅凭表层模式匹配产生偏见。本文提出三项贡献:(i) 构建新任务——基于行为证据的个性推理(GPR),要求模型在评分、推理与证据锚定之间建立链式关联;(ii) 发布新数据集MM-OCEAN(1,104个视频,5,320道多选题),由多智能体流水线生成并经人工验证,包含时间戳行为观察、证据锚定的特质分析及七类线索-证据对应题;(iii) 设计三阶评估体系(评分、推理、证据锚定)及四类样本级失败指标:偏见率(PR)、虚构率(CR)、整合失败率(IR)、整体证据锚定率(HR),对27个MLLM(13个闭源,14个开源)进行评测。结果揭示显著的‘偏见差距’:全领域内51%的正确评分未基于检索到的行为线索,整体证据锚定率仅为0–33.5%。该发现表明,准确的分数不等于正确的推理依据,为实现可解释的社会认知提供了关键路径。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical Big Five score prediction, leaving open whether models truly perceive personality through behavioral understanding or merely prejudge through superficial pattern matching. We address this gap with three contributions. (i) A new task: we formalize Grounded Personality Reasoning (GPR), which requires MLLMs to anchor each Big Five rating in observable evidence through a chain of rating, reasoning, and grounding. (ii) A new dataset: we release MM-OCEAN (1,104 videos, 5,320 MCQs), produced by a multi-agent pipeline with human verification, with timestamped behavioral observations, evidence-grounded trait analyses, and seven categories of cue-grounding MCQs. (iii) Benchmark and analysis: we design a three-tier evaluation (rating, reasoning, grounding) plus four sample-level failure-mode metrics: Prejudice Rate (PR), Confabulation Rate (CR), Integration-failure Rate (IR), and Holistic-grounding Rate (HR), and benchmark 27 MLLMs (13 closed, 14 open). The analysis uncovers a striking Prejudice Gap: across the field, 51% of correct ratings are not grounded in retrieved cues, and the Holistic-Grounding Rate spans only 0-33.5%. These findings expose a disconnect between getting the right score and reasoning for the right reason, charting a roadmap for grounded social cognition in MLLMs.

人格识别多模态偏见检测证据链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。