测试多模态融合中可靠性信息是否真影响决策
When Does Quality-Aware Multimodal Fusion Matter? A Leakage-Safe Diagnostic for Decision-Level Dependence
- 通过打乱可靠性分数检测模型是否真正依赖它
- 实验显示可靠性分数不影响预测,除非能准确预判单模态正确性
- 适合研究多模态融合机制或可信度建模的学者
许多多模态系统会估计各模态的可靠性,并据此加权最终预测。然而,这些可靠性评分是否真正影响模型决策仍不明确。本文提出一种简单诊断方法:训练完成后固定模型和输入,仅对测试样本的可靠性分数进行随机置换。若预测结果依赖于这些分数,性能应下降。在StressID(压力识别)和CMU-MOSEI(情感分析)数据集上的实验表明,即使存在显著提升潜力,置换可靠性分数后性能也保持不变。而在正向对照实验中,当可靠性信号能准确识别正确模态时,相同融合规则带来明显性能提升,说明只有当可靠性信号能可靠预测单模态正确性时,才会影响融合决策。
原文摘要 · Abstract (English)
Many multimodal systems estimate the reliability of each modality and weight their contributions to the final prediction. However, it remains unclear whether these scores influence model decisions or merely correlate with performance. We propose a simple diagnostic to test whether reliability information is used during inference. After training, the model and inputs are fixed while reliability scores are permuted across test examples. If predictions depend on these scores, performance should degrade. Experiments on StressID for stress recognition and CMU-MOSEI for sentiment analysis show that permuting reliability scores leaves performance unchanged despite substantial potential gains from selecting the best modality per example. In positive controls where reliability signals identify the correct modality, the same frozen fusion rules yield significant improvements, indicating that reliability signals influence fused decisions only when they reliably predict unimodal correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。