用2D和3D模型联合训练一个通用视觉编码器,性能接近甚至超越大模型。
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
- 从2D与3D不同任务模型中联合蒸馏,构建统一编码器
- 在分类、分割、3D感知等任务上表现媲美甚至超过原大模型
- 适合需要轻量通用视觉模型的研究与应用
近期多教师蒸馏方法已将多个基础模型的编码器统一为单一编码器,在分类、分割和深度估计等核心视觉任务上表现优异。这引发我们思考:若教师池中包含针对2D与3D感知中多样化任务的专用模型,是否也能取得类似成功?本文首次定义并研究异构教师蒸馏(即共蒸馏)问题,其挑战在于教师模型在设计目标与训练数据上存在显著差异。我们探索了数据共享策略与教师特定编码机制,提出DUNE——一个在2D视觉、3D理解及3D人体感知任务上均表现卓越的单编码器。该模型在各自任务上的性能可媲美其更大的教师模型,甚至在某些情况下超越它们。值得注意的是,DUNE在无地图视觉重定位任务中,以更小的编码器规模超越MASt3R。
原文摘要 · Abstract (English)
Recent multi-teacher distillation methods have unified the encoders of multiple foundation models into a single encoder, achieving competitive performance on core vision tasks like classification, segmentation, and depth estimation. This led us to ask: Could similar success be achieved when the pool of teachers also includes vision models specialized in diverse tasks across both 2D and 3D perception? In this paper, we define and investigate the problem of heterogeneous teacher distillation, or co-distillation, a challenging multi-teacher distillation scenario where teacher models vary significantly in both (a) their design objectives and (b) the data they were trained on. We explore data-sharing strategies and teacher-specific encoding, and introduce DUNE, a single encoder excelling in 2D vision, 3D understanding, and 3D human perception. Our model achieves performance comparable to that of its larger teachers, sometimes even outperforming them, on their respective tasks. Notably, DUNE surpasses MASt3R in Map-free Visual Relocalization with a much smaller encoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。