无需标注数据,让4D重建模型通过自蒸馏自我进化
Self-Improving 4D Perception via Self-Distillation
- 利用时空上下文不对称性设计自蒸馏机制
- 在动态场景中视频深度估计提升36.5%,相机估计提升20.1%
- 适用于多种预训练模型,特别适合缺乏标注的动态场景
大规模多视角重建模型取得了显著进展,但多数方法仍依赖于带真实3D/4D标注的全监督训练。这类标注成本高,尤其在动态场景中极为稀缺,制约了可扩展性。我们提出SelfEvo框架,通过无标签视频持续优化预训练的多视角重建模型。SelfEvo引入基于时空上下文不对称性的自蒸馏方案,使学习型4D感知无需外部标注即可实现自我改进。我们系统研究了提升自改进效果的设计选择,包括损失信号、不对称形式及其他训练策略。在涵盖多个数据集与领域的八个基准上,SelfEvo始终优于预训练基线,并在不同基础模型(如VGGT和$π^3$)间具备良好泛化能力,对动态场景表现尤为显著。整体上,视频深度估计相对提升达36.5%,相机估计提升20.1%,全程未使用任何标注数据。
原文摘要 · Abstract (English)
Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with ground-truth 3D/4D annotations. Such annotations are expensive and particularly scarce for dynamic scenes, limiting scalability. We propose SelfEvo, a self-improving framework that continually improves pretrained multi-view reconstruction models using unlabeled videos. SelfEvo introduces a self-distillation scheme using spatiotemporal context asymmetry, enabling self-improvement for learning-based 4D perception without external annotations. We systematically study design choices that make self-improvement effective, including loss signals, forms of asymmetry, and other training strategies. Across eight benchmarks spanning diverse datasets and domains, SelfEvo consistently improves pretrained baselines and generalizes across base models (e.g. VGGT and $π^3$), with significant gains on dynamic scenes. Overall, SelfEvo achieves up to 36.5% relative improvement in video depth estimation and 20.1% in camera estimation, without using any labeled data. Project Page: https://self-evo.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。