arXiv:2607.27895cs.AIcs.CV2026-07

构建多视角心理理解基准,评估模型对长视频心理状态的深层推理能力。

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

论文配图:MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
图 1 · 摘自论文原文
  • 通过多角色问答生成框架,从第三人称和第一人称双视角设计问题。
  • 包含268段视频、2184个问题,覆盖行为、语境与心理状态的综合推理。
  • 揭示当前大模型在长期视频心理理解上仍存在显著差距,适合心理健康与多模态研究者。

长视频中的心理健康理解需要对可观察行为、人际语境及潜在心理状态进行细致推理。现有基准大多将其简化为粗粒度分类,难以判断模型是否真正理解心理现象,还是依赖表面相关性。为此,我们提出MMHBench,一个全面的多模态心理理解基准,包含268段长视频和2,184个精心设计的问题。评估分为两个互补维度:(1) 第三人称评估,含605个问题,聚焦可观察行为与多模态证据的解释;(2) 第一人称视角共情,含1,579个问题,要求基于多模态证据进行角色化推理。我们提出多智能体问答生成(MAQG)框架,模拟多样社会角色生成问题,经多角色反馈与迭代优化,并由专家验证确保质量。对22个代表性多模态大模型(包括开源与闭源模型)的广泛评估表明,长视频心理健康理解仍极具挑战性。

原文摘要 · Abstract (English)

Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.

心理健康多模态长视频推理评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。