无需训练即可动态调整专家模型,提升医学多模态系统泛化能力。
Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs

- 测试时通过熵引导动态选择专家并在线自适应,无需参数更新。
- 在已见、未见及异构医疗数据上分别提升4.72、7.17、4.3准确率。
- 适合追求零训练成本、强鲁棒性的临床部署场景。
医学视觉语言模型(MVLM)虽具备零样本泛化潜力,但在面对未知模态与领域时可靠性急剧下降,而这正是临床应用最需稳健性的场景。针对此问题,本文从混合专家(MoE)视角重新审视测试时模态泛化,探讨专家是否能在推理阶段无优化条件下实现路由与自适应。我们发现测试时存在根本性专业化与泛化之间的矛盾:盲目聚合专家会稀释模态特异性知识,而仅选单一高置信专家又可能在分布偏移下产生错配。为此,提出完全无需优化的框架MoBE,结合熵引导的动态路由与专家级贝叶斯自适应,使专家可在无梯度更新下在线更新置信度与行为。该方法不改变静态MVLM,仅通过测试时路由与在线统计增强性能,在已见、未见及异构医疗基准上分别取得+4.72、+7.17、+4.3的平均准确率提升,验证了无训练专家自适应对鲁棒模态泛化的有效性。
原文摘要 · Abstract (English)
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where clinical robustness matters most. To address this gap, we revisit test-time modality generalization from the perspective of Mixture-of-Experts (MoE) and ask: can experts route-and-adapt without any optimization during inference? We identify a fundamental specialization-generalization dilemma at test time, where blindly aggregating modality experts dilutes modality-specific knowledge, while selecting one highly confident expert risks mismatch under shift. To address this, we propose MoBE: a fully optimization-free framework that performs dynamic expert selection and adaptation at test time. MoBE combines entropy-guided dynamic routing in MoE settings with expert-wise Bayesian adaptation, enabling experts to update their confidence and adapt online without gradient updates. Without parametric updates, MoBE augments a static MVLM with test-time routing and online statistics, achieving average accuracy gains of +4.72, +7.17, and +4.3 over state-of-the-art TTA methods across seen, unseen, and heterogeneous medical benchmarks, highlighting the effectiveness of training-free expert adaptation for robust modality generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。