用多教师蒸馏训练3D高斯场景编码器,提升泛化能力。
Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- 通过多教师蒸馏,从2D模型中提取语义与结构互补信号。
- 在多项任务上超越基线,仅用39.9倍少的训练场景实现更强性能。
- 适用于3DGS和点云场景,适合需要高效迁移的视觉项目。
尽管3D高斯溅射(3DGS)已成为高保真场景表示方法,但直接从其基本元素中编码丰富通用特征仍缺乏探索。本文提出Chorus,一种多教师预训练框架,通过蒸馏来自语言对齐、通用和物体感知教师的互补信号,学习一个整体的前馈3DGS场景编码器。Chorus采用共享3D编码器与教师专用投影器,促进从高层语义到细粒度结构的统一嵌入空间。我们在多种任务上评估:开放词汇语义与实例分割、线性与解码器探测、数据高效监督及基于大语言模型的问答。此外,我们还测试了仅使用高斯中心、颜色和估计法向量的变体,在仅支持点云的基准上表现出强迁移能力,优于点云基线,且训练场景减少39.9倍。最后,我们提出渲染-蒸馏适配策略,支持域外微调。
原文摘要 · Abstract (English)
While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DGS) scene encoder by distilling complementary signals from 2D foundation models. Chorus employs a shared 3D encoder and teacher-specific projectors to learn from language-aligned, generalist, and object-aware teachers, encouraging a shared embedding space that captures signals from high-level semantics to fine-grained structure. We evaluate Chorus on a wide range of tasks: open-vocabulary semantic and instance segmentation, linear and decoder probing, data-efficient supervision, as well as LLM-based Q&A. Besides 3DGS, we also test Chorus on several benchmarks that only support point clouds by pretraining a variant using only Gaussian centers, colors, and estimated normals. Surprisingly, this encoder shows strong transfer and outperforms the point-cloud baseline while using 39.9 times fewer training scenes. Finally, we propose a render-and-distill adaptation that facilitates out-of-domain finetuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。