arXiv:2608.03475cs.MMcs.AI2026-08

提出可修复不可靠模态的闭环框架,提升多模态意图识别鲁棒性。

Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition

论文配图:Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition
图 1 · 摘自论文原文
  • 通过上下文方差诊断各模态可靠性,结合置信度与跨模态一致性
  • 在噪声、缺失等条件下仍保持准确率,修复后精度显著提升
  • 适合处理实际场景中模态质量不稳的任务,如语音客服、智能助手

多模态意图识别融合语言、语音和视觉信息,但各模态可能噪声大、缺失、语义冲突或主导过强。现有方法通常隐式推断模态重要性,仅重加权或抑制不可靠输入,无法判断退化模态是否可修复并重新信任。本文提出PRIME(Precision-weighted Reliability Inference and Modality rEstoration)框架,实现样本级的可靠性诊断、修复与再评估闭环。该框架通过互补诊断证据(预测置信度、认知分歧、跨模态共识、特征退化)估计每个模态的上下文对数方差。由于缺乏模态可靠性标注,模型在可控模态扰动下训练,并采用异方差不确定性目标。不同于直接丢弃不可靠模态,PRIME利用估计的弱化程度,控制一个原型条件变分修复模块,从其他模态重建退化表示。关键在于修复后重新评估可靠性,决定修复后的表示是否可信参与预测。最终使用修复后的精度进行逆方差多模态融合。在多模态意图识别基准测试中,PRIME在干净数据上表现相当,同时在缺失、噪声、冲突及模态失衡条件下显著提升鲁棒性。

原文摘要 · Abstract (English)

Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant. Existing methods typically infer modality importance implicitly and either reweight or suppress unreliable inputs, without determining whether a degraded modality can be repaired and subsequently trusted. We propose PRIME (Precision-weighted Reliability Inference and Modality rEstoration), a closed-loop reliability guided framework that jointly diagnoses, restores, and reassesses modality quality at the sample level. PRIME represents the weakness of each modality through a contextual log-variance estimated from complementary diagnostic evidence, including predictive confidence, epistemic disagreement, cross-modal consensus, and feature degeneracy. Because modality-reliability annotations are unavailable, the estimator is explicitly trained using controlled modality corruption with known degradation severity, together with a heteroscedastic uncertainty objective. Rather than directly discarding an unreliable modality, PRIME uses its estimated weakness to control a prototype-conditioned variational restoration module that reconstructs the degraded representation from complementary modalities. Crucially, reliability is re-estimated after restoration, allowing the model to determine whether the repaired representation has become sufficiently trustworthy to contribute to prediction. The resulting post-restoration precisions are used for inverse-variance multimodal fusion. Experiments on multimodal intent-recognition benchmarks show that PRIME maintains competitive clean-data performance while improving robustness under missing, noisy, conflicting, and modality-imbalanced conditions.

多模态可靠性诊断模态修复意图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。