arXiv:2410.06558cs.CV2024-10NeurIPS被引 32

通过关联提示设计,让模型在缺失模态时仍能准确识别视觉内容。

Deep Correlated Prompting for Visual Recognition with Missing Modalities

  • 利用提示间的相关性与特征互动,动态生成适应缺失模态的指令。
  • 在三个数据集上均优于现有方法,不同缺失比例下表现稳定。
  • 适合需要鲁棒多模态推理的应用,如隐私敏感的医疗图像分析。

大规模多模态模型在配对多模态数据上表现出色,通常假设输入模态完整。然而现实中因隐私或采集困难常出现模态缺失,导致预训练模型性能下降。为此,我们引入提示学习,将不同缺失情况视为不同输入类型。不同于仅独立拼接提示,我们挖掘提示与输入特征之间的相关性,探索各层提示间的关系,以精细设计指令。同时结合不同模态的互补语义,指导各模态的提示设计。在三个常用数据集上的大量实验表明,本方法在多种缺失场景下均优于先前方法。丰富的消融实验进一步验证了方法在不同缺失比例和类型下的泛化性与可靠性。

原文摘要 · Abstract (English)

Large-scale multimodal models have shown excellent performance over a series of tasks powered by the large corpus of paired multimodal training data. Generally, they are always assumed to receive modality-complete inputs. However, this simple assumption may not always hold in the real world due to privacy constraints or collection difficulty, where models pretrained on modality-complete data easily demonstrate degraded performance on missing-modality cases. To handle this issue, we refer to prompt learning to adapt large pretrained multimodal models to handle missing-modality scenarios by regarding different missing cases as different types of input. Instead of only prepending independent prompts to the intermediate layers, we present to leverage the correlations between prompts and input features and excavate the relationships between different layers of prompts to carefully design the instructions. We also incorporate the complementary semantics of different modalities to guide the prompting design for each modality. Extensive experiments on three commonly-used datasets consistently demonstrate the superiority of our method compared to the previous approaches upon different missing scenarios. Plentiful ablations are further given to show the generalizability and reliability of our method upon different modality-missing ratios and types.

多模态提示学习视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。