arXiv:2510.07355cs.MMcs.SD2025-10被引 1

构建音视频情感推理基准,评估大模型理解情绪并回应的能力

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

  • 设计音视频融合对话数据集,含合成与真实场景
  • 通过感知与交互推理指标衡量模型情绪理解能力
  • 适合研究情感计算与人机交互的学者使用

情感通过语音和面部表情在人机交互中起关键作用。尽管全模态大语言模型发展迅速,但基于音视频线索的情感推理综合评估仍显不足。为此,我们提出 AV-EMO-Reasoning 基准,系统评估大语言模型的情感推理能力。该框架包含经过筛选的音视频语料库,涵盖合成单轮与多轮对话及真实世界子集,并采用情感感知与交互推理指标,检验模型是否能理解用户情绪并生成恰当回应。通过发布可复现的评估标准,该基准为情感感知对话评估提供统一参照,推动更自然、自适应的人机交互发展。

原文摘要 · Abstract (English)

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation of emotional reasoning with audiovisual cues remains limited. To address this gap, we introduce AV EMO Reasoning, a benchmark designed to systematically assess emotional reasoning abilities in large language models. The framework uses a curated audiovisual corpus comprising synthetic single turn and multi turn dialogues and a real world subset, together with emotion perception and interaction reasoning metrics, to evaluate whether models can understand user emotions and produce appropriate responses. By releasing a systematic evaluation benchmark, AV EMO Reasoning offers a reproducible standard for evaluating emotion aware dialogue and advances toward more natural, adaptive human AI interaction.

情感推理音视频大模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。