用DINOv2上采样特征实现无监督目标分割与弱监督材料分割
Upsampling DINOv2 features for unsupervised vision tasks and weakly supervised materials segmentation
- 通过上采样ViT特征,结合聚类实现无监督目标定位与分割
- 在弱监督材料分割任务中表现优异,超越传统方法
- 适合材料表征、属性预测等需要强泛化能力的场景
自监督视觉变换器(ViT)的特征包含对下游任务如目标定位和分割有重要意义的语义与位置信息。近期工作将这些特征与聚类、图分割或区域相关性等传统方法结合,无需微调或训练额外网络即可达到出色基线性能。本文利用ViT网络(如DINOv2)的上采样特征,在两类流程中取得成效:基于聚类的方法用于对象定位与分割;与标准分类器结合用于弱监督材料分割。两项实验均在基准测试中表现良好,尤其在弱监督分割任务中,ViT特征捕捉到经典方法无法获取的复杂关系。我们预期这些特征的灵活性与泛化能力将加速并增强材料表征,涵盖从分割到属性预测的全过程。
原文摘要 · Abstract (English)
The features of self-supervised vision transformers (ViTs) contain strong semantic and positional information relevant to downstream tasks like object localization and segmentation. Recent works combine these features with traditional methods like clustering, graph partitioning or region correlations to achieve impressive baselines without finetuning or training additional networks. We leverage upsampled features from ViT networks (e.g DINOv2) in two workflows: in a clustering based approach for object localization and segmentation, and paired with standard classifiers in weakly supervised materials segmentation. Both show strong performance on benchmarks, especially in weakly supervised segmentation where the ViT features capture complex relationships inaccessible to classical approaches. We expect the flexibility and generalizability of these features will both speed up and strengthen materials characterization, from segmentation to property-prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。