用3D高斯场融合几何与语义信息,提升机器人动作模型的预测能力。
GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

- 通过3D高斯场统一多视角深度、相机参数和语义特征,构建空间对齐的监督信号。
- 在LIBERO-Plus上将FastWAM准确率从52.05%提升至71.29%,Cosmos Policy提升至77.30%。
- 训练后移除所有额外模块,不改变原模型部署结构,适合实际应用。
世界-动作模型(WAMs)联合学习未来视觉预测与动作生成,利用视频动态作为机器人操作的表征学习信号。然而,其视频潜在表示主要优化视觉预测,未显式保留跨视角几何结构或空间局部的对象相关语义。本文提出GaussianWAM,一种训练时的表征增强框架,通过3D高斯场组织几何与语义监督。给定同步多视角观测,冻结的几何与视觉基础模型提供深度、相机参数和密集语义特征。GaussianWAM将这些异构信号绑定到共享的高斯原语,并渲染出空间对齐的语义、深度与覆盖目标,将其蒸馏进WAM的当前观测表征中。所有教师模型、高斯组件及辅助预测头在训练后移除,仅保留原始WAM推理路径,无额外模块或前向计算。在LIBERO-Plus上,GaussianWAM将FastWAM准确率从52.05%提升至71.29%,Cosmos Policy从71.52%提升至77.30%。直接使用CLIP和VGGT蒸馏已建立69.37%的强基线,而3D高斯场统一进一步提升至71.29%,验证了空间组织异构教师信号的优势。GaussianWAM在标准LIBERO上表现更优,并在RoboTwin和真实场景操作中呈现正向迁移趋势。结果表明,训练时的高斯蒸馏为注入几何与语义监督提供了实用方法,且不改变部署架构。
原文摘要 · Abstract (English)
World-Action Models (WAMs) jointly learn future visual prediction and action generation, using video dynamics as a representation-learning signal for robotic manipulation. However, their video latents are primarily optimized for visual prediction and are not explicitly encouraged to preserve cross-view geometric structure or spatially localized, object-relevant semantics. We propose \textbf{GaussianWAM}, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field. Given synchronized multi-view observations, frozen geometry and vision foundation models provide depth, camera parameters, and dense semantic features. GaussianWAM binds these heterogeneous signals to shared Gaussian primitives and renders spatially aligned semantic, depth, and coverage targets, which are distilled into the current-observation representations of the WAM. All teacher models, Gaussian components, and auxiliary prediction heads are removed after training, leaving the original WAM inference path without additional modules or forward computation. On LIBERO-Plus, GaussianWAM improves FastWAM from 52.05\% to 71.29\% and Cosmos Policy from 71.52\% to 77.30\%. Direct CLIP and VGGT distillation already establishes a strong FastWAM baseline of 69.37\%, while Gaussian-field unification further improves it to 71.29\%, supporting the benefit of spatially organizing heterogeneous teacher signals. GaussianWAM also improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation. These results suggest that training-time Gaussian distillation provides a practical way to inject geometry- and semantics-related supervision into WAM representations without changing their deployment architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。