arXiv:2508.10576cs.CV2025-08AAAI被引 10

评测多模态大模型的人类中心交互能力,提升共情与情境理解。

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

  • 构建多模态人类感知基准,评估模型对复杂上下文的理解。
  • 引入渐进式强化学习,显著提升模型在高阶互动任务表现。
  • 发现推理过程具一致思维模式,可无需训练优化非推理模型。

尽管多模态大语言模型(MLLMs)在实现类人交互方面展现出巨大潜力,但进展受限于缺乏针对以人为中心场景的细粒度评估框架,涵盖对复杂人类意图的理解及提供共情、情境感知回应的能力。为此,我们提出HumanSense,一个全面的基准,用于评估MLLM在人类中心感知与交互方面的能力,尤其关注对扩展多模态上下文的深度理解及合理反馈的生成。评估显示,领先模型在高级交互任务上仍有较大提升空间。补充音频与文本信息可显著提升性能,全模态模型在这些任务上表现更优。基于‘恰当反馈源于对对话者需求与情绪的情境分析’的观察,我们提出推理能力是解锁该能力的关键。设计了多阶段、模态渐进的强化学习方法,得到HumanSense-Omni-Reasoning,显著增强高层理解与交互任务表现。此外,我们发现成功的推理过程表现出一致的思维模式,通过设计相应提示,也实现了非推理模型的无训练性能提升。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner.Project page: \textcolor{brightpink}{https://digital-avatar.github.io/ai/HumanSense/}

多模态共情交互推理增强评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。