arXiv:2606.05107cs.CVcs.AI2026-06被引 1

不用标签就能让视觉模型适应科学领域,靠的是已有元数据。

Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

论文配图:Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have
图 1 · 摘自论文原文
  • 用元数据引导自监督学习,替代依赖标签的微调。
  • 在多个科学领域中表现优于无监督与有监督方法。
  • 无需任务标签即可适配模型,适合标签稀缺场景。

我们提出一种无需标签的方法,将通用视觉基础模型适配到特定科学领域。标准监督微调在此类场景中常不适用:标签稀少,且特定任务训练会削弱模型泛化能力并降低鲁棒性。本文改用元数据进行自监督方式的表征适配。所提方法FINO结合标准自监督目标与灵活的元数据引导,可处理细粒度离散和连续元数据,促使表示保留有用因素、抑制无关因素。在亚细胞荧光显微、地球观测、野生动物监测和医学影像等多个领域,FINO均显著优于标准无监督域适应和全监督适配方法,且超越多个领域专用的最先进方法,仅使用无任务标签进行主干适配,监督仅需轻量探针。

原文摘要 · Abstract (English)

We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to these settings: labels are scarce, and task-specific training can collapse the model's generality and hurt robustness. We instead leverage metadata to adapt representations to new domains in a self-supervised manner. Our method, FINO, combines a standard self-supervised objective with flexible metadata guidance that handles both highly granular discrete metadata and continuous metadata. It encourages the representation to preserve informative factors while suppressing spurious ones. Across subcellular fluorescence microscopy, Earth observation, wildlife monitoring, and medical imaging, FINO consistently outperforms standard unsupervised domain adaptation and fully supervised adaptation. It also exceeds highly-specialized domain-specific state of the art, while using no task labels for backbone adaptation and only lightweight probes for supervision.

视觉模型自监督元数据科学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。