arXiv:2606.30597cs.CV2026-06

用稳定潜在提示学习,解决多模态识别中缺失数据问题

Learning from Reliable Latent Prompts for Visual Recognition with Missing Modalities

论文配图:Learning from Reliable Latent Prompts for Visual Recognition with Missing Modalities
图 1 · 摘自论文原文
  • 设计不依赖输入的可学习潜在提示作为稳定锚点
  • 在90%模态缺失下仍保持优异性能,超越现有方法
  • 适合处理真实场景中频繁出现的不完整数据

大规模多模态模型通过融合多样且海量的配对模态信息,在视觉识别任务中表现卓越。然而在真实场景中,缺失模态输入普遍存在,导致为完整模态训练的模型性能急剧下降。现有研究采用提示学习缓解此问题,通常基于实例特征动态生成提示,但该策略在高缺失率下因特征不可靠而失效。本文提出新范式:从可靠潜在提示中学习。我们假设可学习的潜在提示本身蕴含与输入无关的稳定模态先验,因此构建输入无关的可学习提示作为稳定锚点,实现鲁棒引导与有效的跨模态知识补偿,即使在极端缺失率(如90%)下依然有效。在三个基准数据集上的实验表明,所提方法在多种缺失模态场景中均达到当前最优性能,验证了其在应对缺失模态问题上的强大鲁棒性。

原文摘要 · Abstract (English)

Large-scale multimodal models (LMMs) have achieved superior performance in visual recognition by synergizing information across diverse, massive-scale paired modalities. In real-world scenarios, however, missing-modality inputs are ubiquitous, causing models optimized for modality-complete data to exhibit precipitous performance degradation. Existing research has introduced prompt learning to mitigate this issue, typically by generating dynamic prompts from instance-level features, regardless of whether the input modalities are complete or partially absent. However, such input-conditioned strategies are hindered by the escalating unreliability of instance-level features; as higher missing rates increase the proportion of incomplete modalities, the resulting instability in prompt learning limits the model's performance. To address this limitation, we hypothesize that learnable latent prompts themselves encapsulate stable, modality-intrinsic priors that are decoupled from corrupted inputs. Consequently, we propose a novel paradigm: Learning from Reliable Latent Prompts. Unlike prior methods, we model input-agnostic learnable prompts as stable latent anchors that enable robust guidance and effective cross-modal knowledge compensation, even under extreme missing rates (e.g., 90%). Empirical results across three benchmark datasets demonstrate that our "learn-from-latent-prompts" approach achieves state-of-the-art performance across a wide range of missing-modality scenarios. Extensive experiments further confirm the effectiveness of this paradigm in providing a robust solution to the missing-modality problem.

多模态提示学习缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。