聚焦对话情绪热点,提升多模态情感识别精度
Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations
- 定位每句话中情绪显著片段,融合局部与全局特征
- 在多个基准上超越强基线,提升效果稳定可靠
- 适合研究多模态情感分析与跨模态对齐的学者
对话中的情感识别(ERC)困难在于判别性证据稀疏、局部化且各模态间常不同步。本文聚焦情绪热点,提出统一模型:分别检测文本、音频和视频中的每句话情绪热点,通过热点门控融合(HGF)整合局部与全局特征,并使用路由混合对齐器(MoA)实现跨模态对齐;同时构建跨模态图编码对话结构。该设计聚焦显著语段,缓解模态错位,保留上下文信息。在标准ERC数据集上的实验显示,持续优于强基线,消融实验证实了HGF与MoA的有效性。结果表明,以热点为中心的视角可为未来多模态学习提供新思路,推动对话中模态融合的发展。
原文摘要 · Abstract (English)
Emotion Recognition in Conversations (ERC) is hard because discriminative evidence is sparse, localized, and often asynchronous across modalities. We center ERC on emotion hotspots and present a unified model that detects per-utterance hotspots in text, audio, and video, fuses them with global features via Hotspot-Gated Fusion, and aligns modalities using a routed Mixture-of-Aligners; a cross-modal graph encodes conversational structure. This design focuses modeling on salient spans, mitigates misalignment, and preserves context. Experiments on standard ERC benchmarks show consistent gains over strong baselines, with ablations confirming the contributions of HGF and MoA. Our results point to a hotspot-centric view that can inform future multimodal learning, offering a new perspective on modality fusion in ERC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。