DINOv3通过大规模数据与新方法,实现无需标注的通用视觉模型。
DINOv3

- 利用数据优化与新方法提升模型规模与性能。
- 在多任务上超越已有自监督模型,无需微调即达顶尖表现。
- 适合需要高灵活性和跨场景部署的视觉研究者使用。
自监督学习有望消除人工标注需求,使模型能轻松扩展至海量数据和更大架构。通过不针对特定任务或领域,该范式可从自然图像到航拍图像等多种来源中学习视觉表征,仅用单一算法完成。本技术报告介绍 DINOv3,这是迈向这一愿景的重要里程碑,采用简单而有效的策略:首先,通过精心的数据准备、设计与优化,实现数据集与模型规模的协同扩展;其次,提出名为 Gram anchoring 的新方法,有效解决长训练周期中密集特征图退化这一长期未解难题;最后,应用后处理策略进一步提升模型在分辨率、规模及与文本对齐方面的灵活性。结果表明,DINOv3 是一个多功能视觉基础模型,在广泛设置下均优于专用最先进模型,且无需微调。DINOv3 生成高质量密集特征,在多种视觉任务中表现卓越,显著超越此前自监督与弱监督基础模型。我们还发布了 DINOv3 模型套件,旨在推动多样任务与数据上的技术水平,为不同资源约束与部署场景提供可扩展解决方案。
原文摘要 · Abstract (English)
Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this training paradigm has the potential to learn visual representations from diverse sources, ranging from natural to aerial images -- using a single algorithm. This technical report introduces DINOv3, a major milestone toward realizing this vision by leveraging simple yet effective strategies. First, we leverage the benefit of scaling both dataset and model size by careful data preparation, design, and optimization. Second, we introduce a new method called Gram anchoring, which effectively addresses the known yet unsolved issue of dense feature maps degrading during long training schedules. Finally, we apply post-hoc strategies that further enhance our models' flexibility with respect to resolution, model size, and alignment with text. As a result, we present a versatile vision foundation model that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models. We also share the DINOv3 suite of vision models, designed to advance the state of the art on a wide spectrum of tasks and data by providing scalable solutions for diverse resource constraints and deployment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。