arXiv:2608.07183cs.LG2026-08

解决多模态输入缺失时的置信度失准问题,实现自动校准的不确定性估计。

Conformal Fusion Under Missing Modalities

  • 通过模态丢弃训练与证据分解,构建可处理任意模态缺失的融合架构。
  • 在多个真实数据集上对所有模态组合保持目标置信区间覆盖率,覆盖误差显著降低。
  • 适合需要可靠不确定性估计的机器人、医疗等高风险多模态场景使用。

多模态融合模型通常假设推理时所有模态均可用,但传感器故障、采集差异和成本限制常导致观测不完整。现有方法将模态缺失视为精度问题,却忽略了置信度是否仍校准这一根本问题。本文提出模态条件化置信融合(MCCF),同时解决缺失鲁棒性与校准不确定性。MCCF结合多模态瓶颈融合主干(训练时加入模态丢弃)、各模态独立的证据头生成模态解耦的狄利克雷分布,以及基于德普斯特-沙弗规则的证据融合机制;缺失模态贡献空证据,结构上被忽略,因此融合后的不确定性自然反映信息减少,无需测试时补全。再通过基于模态存在掩码的蒙德里安置信校准模块,为每个非空模态子集提供有限样本下的组条件覆盖率。MCCF是首个通过架构集成实现任意模态可用性下形式化覆盖率保证的方法,而非事后校准;其证据分解还能生成各模态的空缺得分,定位不确定性的来源。在合成任务和三个真实多模态基准上,MCCF在所有模态存在子集上均保持目标覆盖率,相对于边际分割置信基线显著缩小了全模态与部分模态间的覆盖率差距,且相比温度缩放和证据基线,未带来可测量的精度损失。

原文摘要 · Abstract (English)

Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality absence as a prediction-accuracy problem, leaving a more basic question unanswered: whether a model's confidence estimates remain calibrated when an entire input stream is removed. We argue that missing-modality robustness and calibrated uncertainty are a single coupled property, and introduce Modality-Conditioned Conformal Fusion (MCCF), an architecture that addresses both at once. MCCF combines a multimodal bottleneck fusion backbone trained with modality dropout, per-modality evidential heads producing modality-decomposed Dirichlet distributions, and a Dempster-Shafer combination rule that fuses the per-modality evidence into a joint predictive distribution; an absent modality contributes vacuous evidence that is structurally ignored, so the fused uncertainty automatically reflects the reduced information without test-time imputation. A Mondrian conformal calibration module keyed on the modality-presence mask then provides finite-sample group-conditional coverage for every non-empty modality subset. MCCF is, to our knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty to the absent modality responsible. Across a synthetic problem and three real multimodal benchmarks, MCCF holds its target coverage on every modality-presence subset, substantially narrows the coverage gap between full and partial modalities relative to a marginal split-conformal baseline, and imposes no measurable accuracy cost relative to temperature-scaled and evidential baselines.

多模态不确定性置信度校准鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。