将多视角视觉知识提炼到单视角模型,提升3D一致性与局部辨识性。
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models

- 用多视角几何特征融合2D基础特征,构建判别性教师模型。
- 在保持语义结构前提下,使特征兼具3D一致性和局部区分性。
- 适用于需要精准对应关系的3D视觉任务,如多视图重建与渲染。
基础视觉特征(如DINO)在现代计算机视觉中扮演关键角色,近年来也成为多视角前馈几何估计的核心组件。本文证明,通过将这些多视角模型内部蕴含的3D几何知识重新蒸馏到单视角估计器中,可获得更优的3D一致基础特征。核心思路是构建一个融合预训练2D基础特征与多视角几何特征的多视角教师模型,并利用判别性排序目标优化融合表示。所提出的判别蒸馏框架使学习到的特征既具备3D一致性,又具有局部独特性,同时保持与原始基础模型特征空间对齐,以保留预训练表征的语义结构。一致性与局部可区分性对3D视觉任务(如图像间语义与几何对应)至关重要。通过涵盖直接特征分析、密集预测迁移及显式3D提升与渲染的全面实验,验证了该方法的有效性:所生成的特征显著提升多视图一致性与局部可区分性,同时保持原始表征的语义可迁移性。
原文摘要 · Abstract (English)
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。