解决视觉语言模型部署时模态偏移问题,提升预测可靠性。
Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift

- 通过约束解混合机制,仅优化轻量级融合模块。
- 文本或联合偏移下准确率提升至66.51%和26.27%。
- 适合应对模态可靠性不一致的现实场景。
视觉-语言模型在部署时可能遭遇不对称的视觉与文本漂移。此类漂移暴露了多模态失效模式:不可靠分支仍保持高置信度,主导融合过程,导致基于熵的测试时自适应反而强化错误预测。本文将此行为建模为双重随机后验混合,并将自适应视为受约束的解混合问题。提出的主序引导多模态测试时自适应(MG-MTTA)冻结两个编码器,仅更新轻量级融合模块。运行锚点一致性估计各分支漂移程度,跨模态冲突在熵锐化前调节模态主导性。分析给出了熵降低保留正确决策的充分条件,以及偏差模态反转融合排序的显式阈值。在视觉、文本及联合漂移下,MG-MTTA使ImageNet top-1准确率在文本漂移下从57.97%提升至66.51%,在联合漂移下从21.68%提升至26.27%,同时减少错误高置信度失败。最大增益出现在文本与联合漂移场景,此时两分支可靠性差异更大。
原文摘要 · Abstract (English)
Vision--language models can face asymmetric visual and textual shifts at deployment. These shifts expose a multimodal failure mode in which an unreliable branch remains overconfident, dominates fusion, and causes entropy-based test-time adaptation to sharpen an incorrect prediction. We model this behavior as doubly stochastic posterior mixing and cast adaptation as constrained de-mixing. Majorization-Guided Multimodal Test-Time Adaptation (MG-MTTA) freezes both encoders and updates only a lightweight fusion module. Running-anchor consistency estimates relative branch drift, while cross-modal conflict regulates modality dominance before entropy sharpening. The analysis gives sufficient conditions for entropy reduction to preserve the clean decision and an explicit threshold at which a biased modality reverses the fused ranking. Across visual, textual, and joint shifts, MG-MTTA improves ImageNet top-1 accuracy from 57.97\% to 66.51\% under textual shift and from 21.68\% to 26.27\% under joint shift, while reducing wrong-more-confident failures. The largest gains occur under textual and joint shifts, where the two branches differ more in reliability. Project page: https://mg-mtta.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。