提出可自适应决定是否修复缺失模态的SIEVE模型
Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?
- 基于样本级充分性信号,动态判断是否需修复缺失模态
- 在CMU-MOSI和IEMOCAP上优于主流修复方法,接近最优性能
- 无需改动现有修复模块,可即插即用,适合实际部署
现有多模态情感分析方法在模态缺失时通常采用先修复后预测的范式。我们重新审视这一假设,发现并非所有样本都需修复全部模态:仅少数样本需要全模态输入,而不同模态子集对不同样本更优。基于此,我们提出SIEVE模型,将‘是否修复’转化为样本级可学习决策。SIEVE通过比较直接预测分支与修复分支的损失差距,生成经验充分性信号,并利用证据门控机制联合建模模态充分性及其认知不确定性。该方法不依赖特定修复模块,可作为通用决策层嵌入任意修复框架。在CMU-MOSI和IEMOCAP数据集上的实验表明,SIEVE在不同缺失率下均显著提升主流修复模型性能,逼近每样本双分支的理论最优解。
原文摘要 · Abstract (English)
Existing methods for multimodal sentiment analysis (MSA) under missing modalities usually follow a repair-first paradigm. We revisit this assumption and ask: \emph{should every missing modality be repaired?} A per-sample oracle analysis shows the answer is not always: full-modality input is optimal for only a small fraction of samples, and every modality subset is preferred by some samples. These results suggest that adding or repairing modalities may not always improve prediction, and that the utility of each modality is sample-dependent. Building on this finding, we propose \textbf{S}ufficiency-\textbf{I}nformed \textbf{E}vidential \textbf{V}al\textbf{vE} (\textbf{SIEVE}) that turns ``whether to repair'' into an explicit, learnable decision at the sample level. SIEVE compares a direct prediction branch with a repair branch, derives an empirical sufficiency signal from their per-sample loss gap, and routes each input through an evidential gate that jointly models sufficiency and its epistemic uncertainty. SIEVE is repair-agnostic: it operates as a plug-and-play decision on top of any explicit or implicit repair module, without modifying its internal design. Experiments on CMU-MOSI and IEMOCAP show that SIEVE consistently improves representative repair backbones across evaluated missing rates, and approaches the per-sample dual-branch achievable optimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。