arXiv:2505.14204cs.CVq-bio.NC2025-05

用人类感知初始化视觉模型,让AI更懂人眼所见

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment

  • 用人类感知数据初始化CLIP视觉编码器,而非后期微调
  • 29个零样本任务中,图像分类与检索性能显著提升
  • 无需领域适配,适合构建通用视觉语言系统

我们提出感知初始化(Perceptual-Initialization, PI),将人类感知结构融入视觉表征学习的初始化阶段,而非作为下游微调步骤。通过使用来自NIGHTS数据集的人类衍生三元组嵌入初始化CLIP视觉编码器,并在YFCC15M上进行自监督预训练,该方法在29个零样本分类和检索基准上均实现显著提升,且无需任何任务特定微调。在ImageNet-1K上,约经过15个训练周期后即出现零样本性能增益。不同规模数据集均表现出改进效果,且提升在预训练过程中的不同阶段显现,具体取决于数据特性。该方法持续提升零样本准确率(如top-1、top-5)和检索召回率(如R@1、R@5),且无需针对目标领域进行调整。结果挑战了将人类感知数据仅用于微调的传统观念,表明在早期表征学习中嵌入人类感知结构可构建更具泛化能力、更对齐的视觉语言系统。本工作证明‘始于你’——从人类感知出发——为通用视觉语言智能奠定更强基础。

原文摘要 · Abstract (English)

We introduce Perceptual-Initialization (PI), a paradigm shift in visual representation learning that incorporates human perceptual structure during the initialization phase rather than as a downstream fine-tuning step. By integrating human-derived triplet embeddings from the NIGHTS dataset to initialize a CLIP vision encoder, followed by self-supervised learning on YFCC15M, our approach demonstrates significant zero-shot performance improvements, without any task-specific fine-tuning, across 29 zero shot classification and 2 retrieval benchmarks. On ImageNet-1K, zero-shot gains emerge after approximately 15 epochs of pretraining. Benefits are observed across datasets of various scales, with improvements manifesting at different stages of the pretraining process depending on dataset characteristics. Our approach consistently enhances zero-shot top-1 accuracy, top-5 accuracy, and retrieval recall (e.g., R@1, R@5) across these diverse evaluation tasks, without requiring any adaptation to target domains. These findings challenge the conventional wisdom of using human-perceptual data primarily for fine-tuning and demonstrate that embedding human perceptual structure during early representation learning yields more capable and vision-language aligned systems that generalize immediately to unseen tasks. Our work shows that "beginning with you", starting with human perception, provides a stronger foundation for general-purpose vision-language intelligence.

视觉语言感知初始化零样本表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。