arXiv:2606.31329cs.ROcs.AI2026-06

让机器人规划更准:3D轨迹直接指导操作,避免2D误导

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

论文配图:3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
图 1 · 摘自论文原文
  • 用3D深度重建增强视觉语言模型,直接生成三维路径点
  • 在仿真和真实场景中表现优于传统2D引导方法,尤其在环境变化时
  • 适合做复杂场景下机器人自主操作的研究者和开发者

分层视觉-语言-动作模型通过解耦高层规划与底层控制提升机器人操作的泛化能力。现有方法使用视觉语言模型预测2D末端执行器轨迹作为下游策略的显式引导,但顶尖低层策略基于点云在3D度量空间运行,输入缺乏深度信息的2D轨迹会导致每个航点被赋予其下方场景表面的深度,产生几何失真。我们提出3D HAMSTER,一种通过规划器直接输出度量可靠的3D轨迹来弥合这一差距的分层框架。通过为视觉语言模型增加专用深度编码器和密集深度重建目标,实现3D航点序列的预测,并直接集成到基于点云的低层策略中。在3D轨迹预测、仿真和真实世界操作任务中,3D HAMSTER始终优于专有视觉语言模型和2D引导基线,尤其在外观改变、未知语言、空间和视觉条件下提升显著。

原文摘要 · Abstract (English)

Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.

机器人操作3D轨迹视觉语言模型分层控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。