通过迭代修正代理增强不完整多模态情感分析的鲁棒性
Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

- 用非语言模态构建语言代理,并在多模态上下文中逐步优化
- 在多种缺失设置下,比基线模型提升3.2%~5.1%准确率
- 适合处理真实场景中数据缺失的情感分析任务
多模态情感分析旨在融合语言、视觉和声学线索推断情感状态。然而,现实中的多模态输入常存在缺失或损坏,削弱跨模态互补性并引入误导信息。现有基于代理的方法通常采用单次构建代理来补偿退化的语言信息,但初始代理可能粗糙或不可靠。过早注入此类代理会传播初始误差,影响情感预测。为此,我们提出一种迭代代理修正框架以提升不完整多模态情感分析的鲁棒性。该方法从非语言模态构建语言导向代理,并通过门控残差修正在多模态上下文中逐步优化。经修正的代理根据语言可靠性评分自适应融合观测到的语言表征,实现代理补偿与可信语言证据间的平衡。此外,引入分阶段潜在修正目标,利用完整语言表示作为训练期语义锚点,稳定代理优化轨迹。在MOSI、MOSEI和SIMS数据集上,多种缺失设置下的实验表明,该框架持续优于竞争基线,在不完整输入下实现稳健的情感预测。
原文摘要 · Abstract (English)
Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。