用3D高斯点云重建新视角,提升机器人策略的泛化能力
Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
- 基于3D高斯溅射的单次前向重建,从稀疏无标定输入恢复高质量3D场景
- 通过3D先验蒸馏防止几何坍塌,生成稳定3D表示并渲染多样化合成视角
- 在严重空间扰动下性能远超基线,适合需要强视角泛化的机器人控制任务
当前主流的2D中心视觉-运动策略在新视角泛化方面存在明显缺陷,因其依赖静态观测,难以在未见视角间保持一致的动作映射。为此,我们提出GenSplat,一种前馈式3D高斯溅射框架,通过新视角渲染实现视图泛化策略学习。GenSplat采用置换等变架构,在一次前向传播中即可从稀疏、未标定输入重建高保真3D场景。为保证结构完整性,设计了3D先验蒸馏策略,对3DGS优化进行正则化,防止仅依赖光度监督导致的几何坍塌。通过从这些稳定的3D表示中渲染多样化的合成视角,系统性地扩充训练过程中的观测流形。该增强迫使策略基于底层3D结构作出决策,从而在严重空间扰动下仍能稳健执行,而基线方法性能显著下降。
原文摘要 · Abstract (English)
Prevailing 2D-centric visuomotor policies exhibit a pronounced deficiency in novel view generalization, as their reliance on static observations hinders consistent action mapping across unseen views. In response, we introduce GenSplat, a feed-forward 3D Gaussian Splatting framework that facilitates view-generalized policy learning through novel view rendering. GenSplat employs a permutation-equivariant architecture to reconstruct high-fidelity 3D scenes from sparse, uncalibrated inputs in a single forward pass. To ensure structural integrity, we design a 3D-prior distillation strategy that regularizes the 3DGS optimization, preventing the geometric collapse typical of purely photometric supervision. By rendering diverse synthetic views from these stable 3D representations, we systematically augment the observational manifold during training. This augmentation forces the policy to ground its decisions in underlying 3D structures, thereby ensuring robust execution under severe spatial perturbations where baselines severely degrade.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。