arXiv:2507.07802cs.CV2025-07ICCV被引 13

动态提示融合策略提升视觉识别在缺失模态下的鲁棒性

Synergistic Prompting for Robust Visual Recognition with Missing Modalities

  • 设计动态适配器,根据缺失情况自适应生成提示
  • 静态与动态提示协同,使关键信息缺失时仍保持高精度
  • 在三个数据集上验证,对不同缺失率均表现稳定

大规模多模态模型在丰富配对数据训练下,在各类视觉识别任务中表现出色。然而在实际应用中,模态缺失或不完整常导致性能显著下降。现有基于提示的方法存在两大局限:(1) 静态提示难以适应不同缺失场景;(2) 基础提示调优在关键模态缺失时可靠性不足。为此,我们提出一种新型协同提示(Synergistic Prompting, SyP)框架,以应对多模态缺失问题。该框架包含两项核心创新:(I) 动态适配器,通过计算自适应缩放因子,动态生成提示,替代静态参数,实现灵活的多模态适配;(II) 协同提示策略,融合静态与动态提示,平衡模态间信息,确保关键模态缺失时仍具备鲁棒推理能力。在三个主流视觉识别数据集上的实验表明,该方法在多种缺失率和条件下均显著优于现有方法。大量实验与消融分析验证了其在处理缺失模态方面的有效性,凸显其卓越的适应性与可靠性。

原文摘要 · Abstract (English)

Large-scale multi-modal models have demonstrated remarkable performance across various visual recognition tasks by leveraging extensive paired multi-modal training data. However, in real-world applications, the presence of missing or incomplete modality inputs often leads to significant performance degradation. Recent research has focused on prompt-based strategies to tackle this issue; however, existing methods are hindered by two major limitations: (1) static prompts lack the flexibility to adapt to varying missing-data conditions, and (2) basic prompt-tuning methods struggle to ensure reliable performance when critical modalities are missing.To address these challenges, we propose a novel Synergistic Prompting (SyP) framework for robust visual recognition with missing modalities. The proposed SyP introduces two key innovations: (I) a Dynamic Adapter, which computes adaptive scaling factors to dynamically generate prompts, replacing static parameters for flexible multi-modal adaptation, and (II) a Synergistic Prompting Strategy, which combines static and dynamic prompts to balance information across modalities, ensuring robust reasoning even when key modalities are missing. The proposed SyP achieves significant performance improvements over existing approaches across three widely-used visual recognition datasets, demonstrating robustness under diverse missing rates and conditions. Extensive experiments and ablation studies validate its effectiveness in handling missing modalities, highlighting its superior adaptability and reliability.

多模态提示学习鲁棒性缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。