评估多模态医疗模型在传感器缺失时的鲁棒性,为临床AI系统选型提供依据。
MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion

- 构建涵盖9个临床数据集、6种融合架构的评测基准,模拟两种缺失场景。
- 模型家族比参数量更能预测鲁棒性,短序列对内部缺失更敏感。
- 适合临床AI研发者与多模态融合方法设计者参考,尤其关注传感器失效场景。
多模态生理数据支撑从重症监护到可穿戴设备的临床AI系统,但传感器故障在实际中常见。主要故障模式包括:模态缺失(整条通道丢失)和模态内缺失(连续时间段丢失)。现有基准无法在多种融合架构下,于不同严重度级别上,跨多样临床数据集评估这两种故障模式。本文提出MuteBench,覆盖7个临床领域中的9个数据集、6种融合架构,以及2种缺失模式,共12.5万样本。结果表明,模型家族是鲁棒性的最强预测因子,超越参数量;独立通道模型虽能较好容忍模态缺失,但在模态内缺失(尤其短序列)下表现脆弱。课程式模态丢弃仅在训练中最大丢弃率范围内有效。此外,通道数、序列长度和模态对齐共同决定哪种故障模式威胁更大。在PTB-XL数据集上的案例研究显示,基于扩散的插补可提升下游分类性能,尤其对专家路由机制敏感的模型增益最大,但跨数据集验证仍待深入。MuteBench为从业者提供了选择现有架构及设计未来鲁棒融合方法的具体指导。
原文摘要 · Abstract (English)
Multimodal physiological data powers clinical AI systems from intensive care units to wearable devices, but sensors routinely fail in practice. Two failure modes are common: modality missing, where an entire channel is absent, and within-modality missing, where a contiguous time segment is lost. No existing benchmark evaluates multiple fusion architectures under both failure modes at controlled severity levels across diverse clinical datasets. We present MuteBench, a benchmark covering 9 datasets from 7 clinical domains, 6 fusion architectures, and 2 missing-data modes over 125,000 samples. Through this benchmark, we find that architecture family is the strongest predictor of robustness, outweighing parameter count. Channel-independent models tolerate modality missing well but can be sensitive to within-modality missing, especially on short sequences. Curriculum modality dropout protects reliably only up to the maximum dropout rate used in training. We also find that channel count, sequence length, and modality alignment jointly determine which failure mode poses the greater threat. Finally, a PTB-XL case study suggests that diffusion-based imputation can improve downstream classification under within-modality missing, with the largest gains for models whose expert routing is most sensitive to corrupted inputs, though broader validation across datasets remains an open direction. MuteBench provides practitioners with concrete guidance for both selecting existing architectures and informing the design of future robust multimodal fusion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。