用偏好优化提升多模态情绪理解,减少错误关联和幻觉。
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
- 通过构建偏好数据,让模型更关注真实音视频线索。
- 在多个数据集上实现6%-19%的零样本性能提升。
- 适合研究社交智能、多模态推理与模型对齐的学者。
情绪理解是构建社会智能体的关键。尽管近期多模态大语言模型在此任务上表现强劲,但仍面临两大挑战:情绪与无关音视频线索之间的虚假关联,以及由语言模型文本先验引发的音视频线索幻觉。为此,我们提出EmoReAlM基准,用于评估模型在线索-情绪关联、幻觉和模态一致性方面的能力。进一步提出AVEm-DPO方法,通过偏好优化使模型响应更契合音视频输入与以情绪为中心的查询。具体地,构建包含虚假关联或幻觉响应的偏好对,并引入由文本提示引导的音视频输入对;同时加入正则项惩罚对文本先验的依赖,缓解特定模态线索的幻觉问题。在DFEW、RAVDESS和EMER上的实验表明,该方法在零样本设置下相较基线模型实现6%-19%的相对性能提升。本工作提供了严谨的评估基准与稳健的优化框架,推动多模态大模型在情绪理解与社会智能方向的发展。代码、模型与基准将开源于https://avere-iclr.github.io。
原文摘要 · Abstract (English)
Emotion understanding is essential for building socially intelligent agents. Although recent multimodal large language models have shown strong performance on this task, two key challenges remain - spurious associations between emotions and irrelevant audiovisual cues, and hallucinations of audiovisual cues driven by text priors in the language model backbone. To quantify and understand these issues, we introduce EmoReAlM, a benchmark designed to evaluate MLLMs for cue-emotion associations, hallucinations and modality agreement. We then propose AVEm-DPO, a preference optimization technique that aligns model responses with both audiovisual inputs and emotion-centric queries. Specifically, we construct preferences over responses exhibiting spurious associations or hallucinations, and audiovisual input pairs guided by textual prompts. We also include a regularization term that penalizes reliance on text priors, thereby mitigating modality-specific cue hallucinations. Experimental results on DFEW, RAVDESS and EMER demonstrate that our method significantly improves the performance of the reference baseline models with 6-19% of relative performance gains in zero-shot settings. By providing both a rigorous benchmark and a robust optimization framework, this work enables principled evaluation and improvement of MLLMs for emotion understanding and social AI. Code, models and benchmark will be released at https://avere-iclr.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。