让缺失模态也能融合,用代理填充保持接口一致
COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing
- 用代理表示填充缺失模态,恢复完整输入结构
- 在多个数据集上提升多种缺失场景下的鲁棒性
- 适合实际部署中模态不全的感知系统
多模态感知中缺失模态不仅导致信息丢失,还引发融合接口不匹配:训练时固定的模态槽位在推理时面对变化的观测子集。我们提出 Compass,一种接口完备的融合框架,在预测前恢复标准槽位结构。每个模态分配固定融合槽位,已观测模态用真实表示填充,缺失模态则由来自可观测源估计的靶向槽位补全表示填充。同一缺失槽位的多个源估计被聚合为单一填充项,使轻量级融合算子可适配任意缺失模式。训练采用合成模态掩码、槽位兼容性监督与表示空间稳定化,确保补全槽位与真实表示兼容且利于下游识别。在 XRF55、MM-Fi 与 OctoNet 上,Compass 在单模态与多模态缺失设置下均表现更优,优于插补、蒸馏与转换类基线。结果表明,保持融合接口是实现鲁棒多模态感知的简洁有效原则。
原文摘要 · Abstract (English)
Missing modalities in multimodal sensing cause not only information loss but also a fusion-interface mismatch: a fusion head trained on a canonical set of modality slots must operate on changing observed subsets at inference time. We propose Compass, an interface-complete fusion framework that restores this canonical slot structure before prediction. Each modality is assigned a fixed fusion slot. Observed modalities populate their slots with real representations, while absent modalities are filled with target-slot completion representations estimated from the observed sources. Multiple source-specific estimates for the same missing slot are aggregated into a single slot filler, allowing the same lightweight fusion operator to be applied under arbitrary missing-modality patterns. Training uses synthetic modality masking, slot-compatibility supervision, and representation-space stabilization to make completed slots compatible with real modality representations and useful for downstream recognition. Across XRF55, MM-Fi, and OctoNet, Compass improves robustness under diverse single- and multiple-missing settings, including controlled comparisons against imputation, distillation, and translation-style baselines. These results suggest that preserving the fusion interface is a simple and effective principle for robust multimodal sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。