arXiv:2411.15099cs.CVcs.CL2024-11CVPR被引 4

让视觉语言模型更擅长少样本适应,同时保持零样本能力。

Context-Aware Multimodal Pretraining

  • 引入上下文感知的预训练目标,增强模型对额外信息的利用能力。
  • 在21个下游任务中,少样本效率最高提升4倍,平均性能增益超5%。
  • 无需训练即可通过简单度量实现优越泛化,适合快速迁移场景。

大规模多模态表征学习在测试时实现了良好的零样本迁移能力。然而,标准预训练范式(基于大量图文数据的对比学习)并未显式促进表征支持少样本适应。本文提出一种简单但精心设计的多模态预训练扩展,使表征能有效利用额外上下文。实验表明,使用该目标训练的视觉-语言模型显著提升了少样本适应能力:在21个下游任务中,测试时样本效率最高提升四倍,平均少样本适应性能提升超过5%,且在不同模型规模和训练时长下均保持零样本泛化性能。特别地,仅需无需训练的基于度量的简单适应机制,其表现即超越更复杂、成本更高的优化方法,极大简化了新领域泛化过程。

原文摘要 · Abstract (English)

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amounts of image-text data) does not explicitly encourage representations to support few-shot adaptation. In this work, we propose a simple, but carefully designed extension to multimodal pretraining which enables representations to accommodate additional context. Using this objective, we show that vision-language models can be trained to exhibit significantly increased few-shot adaptation: across 21 downstream tasks, we find up to four-fold improvements in test-time sample efficiency, and average few-shot adaptation gains of over 5%, while retaining zero-shot generalization performance across model scales and training durations. In particular, equipped with simple, training-free, metric-based adaptation mechanisms, our representations easily surpass more complex and expensive optimization-based schemes, vastly simplifying generalization to new domains.

多模态学习少样本适应预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。