arXiv:2605.16889cs.CV2026-05中稿 · IJCAI被引 1

解决多模态情感分析中缺失模态导致的预测漂移问题

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

论文配图:Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
图 1 · 摘自论文原文
  • 通过两级参考对齐,稳定不同模态组合的特征表示
  • 在多种缺失模式下保持稳定预测,全模态时准确率超86%
  • 适合处理真实场景中不完整或质量差的多模态数据

多模态情感分析依赖文本、语音和视觉信号,但现实数据常存在模态缺失和质量失衡。现有方法从可用模态生成缺失模态特征,但因各模态表达机制和情感动态差异,生成特征可能偏离真实分布,误导预测。此外,不可靠模态可能主导融合,导致不同模态组合间表示偏移,产生不稳定的情感表征。为此,我们提出两级参考对齐框架:第一级利用完整模态样本约束表示,将不同模态组合对齐至共享情感空间;第二级通过原型检索与投票抑制不可靠模态,在决策层实现跨模态一致性。实验表明,该框架在CMU-MOSI和CMU-MOSEI数据集上,于多种缺失设置下均表现稳定。全模态输入时,准确率分别达86.28%和85.88%,F1值分别为86.24%和85.86%,达到当前最优水平。

原文摘要 · Abstract (English)

Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods generate features for modality missing from available ones, but differences in expression mechanisms and sentiment dynamics across modalities may cause the generated features to deviate from true distributions and mislead prediction. In addition, unreliable modalities may dominate fusion, resulting in representation shift across modality combinations and unstable sentiment representations. To address these challenges, we propose a two-level reference alignment framework. The framework introduces stable references at the feature representation and sentiment decision levels to improve robustness under modality missing. First-level reference alignment leverages complete-modality samples to constrain representations and align different modality combinations into a shared sentiment space. Second-level reference alignment enforces cross-modal consistency at the decision level by suppressing unreliable modalities through prototype retrieval and voting. As a result, the framework maintains stable and reliable sentiment predictions under diverse missing-modality patterns. Experiments on CMU-MOSI and CMU-MOSEI show consistent improvements across various missing-modality settings. Under full-modality input, the proposed method achieves state-of-the-art performance, with ACC of 86.28% and 85.88%, and F1 of 86.24% and 85.86%.

多模态情感分析缺失模态稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。