分离模态处理与融合,提升情感分析准确率
Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

- 分三步:先分离模态、再分别优化、最后融合
- 在两个数据集上均刷新最高性能,全面超越现有方法
- 无需额外监督,能自动调整模态权重应对干扰
多模态融合需同时优化模态内信号和跨模态交互,但二者常混杂于同一过程。本文提出SeRIn(Segregate, Refine, Integrate)架构,将此双重目标作为结构先验强制分离:各模态沿独立路径演化,分别在对应编码器上下文中进行细化;专用跨模态路径积累联合演化信息而不污染单模态流。完整的跨模态交互延后至最终预测阶段——消融实验证明,性能提升源于结构化交互而非容量增加;在视觉损坏条件下,门控分析显示模型自发实现模态重加权,无须显式监督。SeRIn在CH-SIMS和CMU-MOSEI两个基准上均达到当前最优,各项指标全面领先。
原文摘要 · Abstract (English)
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。