arXiv:2605.19950cs.CV2026-05被引 3

让模型预测情绪变化,提升多模态情感计算的动态理解能力

AffectVerse: Emotional World Models for Multimodal Affective Computing

论文配图:AffectVerse: Emotional World Models for Multimodal Affective Computing
图 1 · 摘自论文原文
  • 用跨模态时序想象预测未来情绪表示,实现短期情绪演变建模
  • 在9个基准上性能提升至少2.57%,证明预测性信念建模有效
  • 适合研究情绪动态建模、人机情感交互的开发者与研究人员

人类通过整合多模态线索与对情感状态演变的预期来推断情绪。现有多模态大语言模型(MLLM)通常将情绪识别视为对完整音视频文本输入的静态融合,忽略了情感动态。我们提出AffectVerse,基于Qwen2.5-Omni,配备情绪世界模块(EWM),一个无需动作、仅在表征层进行短时程潜情绪预测的模块。EWM包含三个组件:1)跨模态时序想象,从历史标记中多步滚动预测未来视频/音频表示;2)模态感知多步注意力(MAMA)信念聚合,将想象标记压缩为模态感知的信念标记;3)信念注入,将这些信念标记插入大语言模型进行情感推理。AffectVerse以未来预测作为过去条件的自监督信号:不替代历史建模,也不依赖推理时的未见信号,而是强制当前信念状态编码可预测后续情感变化的转移线索。在9个基准上,AffectVerse性能优于其他模型至少2.57%。控制消融实验表明,时序想象、跨模态滚动和信念聚合均带来增量增益。结果表明,预测性信念状态建模是情感计算的可行替代方案。

原文摘要 · Abstract (English)

Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static fusion over complete audiovisual-text inputs, leaving affective dynamics implicit. We propose AffectVerse, a Qwen2.5-Omni-based model equipped with an Emotion World Module (EWM), an action-free representation-level module for short-horizon latent affective prediction. \rev{EWM contains three modules: 1) Cross-Modal Temporal Imagination predicts future video/audio representations from past tokens with multi-step rollout. 2) MAMA(Modality-Aware Multi-step Attention) Belief Aggregation compresses imagined tokens into modality-aware belief tokens. 3) Belief Injection inserts these belief tokens into the LLM for affective reasoning.} AffectVerse uses future prediction as a past-conditioned self-supervised signal: it does not replace modeling observed history or require unseen signals at inference, but forces the current belief state to encode transition cues that are predictive of subsequent affective change. Across nine benchmarks, AffectVerse improves at least 2.57\% over other models, while controlled ablations show additive gains from temporal imagination, cross-modal rollout, and belief aggregation. These results suggest predictive belief-state modeling is a practical alternative for affective computing.

情感计算多模态情绪预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。