arXiv:2603.17993cs.CVcs.RO2026-03被引 1

GMT通过融合多模态信息,生成6自由度物体在3D场景中的精准可控轨迹。

GMT: Goal-Conditioned Multimodal Transformer for 6-DOF Object Trajectory Synthesis in 3D Scenes

  • 采用多模态变压器联合建模几何、语义与目标姿态信息
  • 在合成与真实数据上显著提升空间精度与朝向控制能力
  • 适合需要复杂3D交互的机器人路径规划研究者

在3D环境中合成可控制的6自由度物体操作轨迹对机器人交互至关重要,但受限于精确的空间推理、物理可行性及多模态场景理解。现有方法多依赖2D或部分3D表示,难以捕捉完整场景几何,限制轨迹精度。我们提出GMT,一种多模态变压器框架,通过联合利用3D边界框几何、点云上下文、语义物体类别和目标末端位姿,生成现实且目标导向的物体轨迹。模型将轨迹表示为连续的6-DOF姿态序列,并采用定制化条件策略融合几何、语义、上下文与目标导向信息。在合成与真实世界基准上的大量实验表明,GMT优于当前先进的人体动作与人-物交互基线(如CHOIS和GIMO),在空间准确性和朝向控制方面取得显著提升。该方法建立了基于学习的操纵规划新基准,并在多样化物体与杂乱3D环境展现出强泛化能力。

原文摘要 · Abstract (English)

Synthesizing controllable 6-DOF object manipulation trajectories in 3D environments is essential for enabling robots to interact with complex scenes, yet remains challenging due to the need for accurate spatial reasoning, physical feasibility, and multimodal scene understanding. Existing approaches often rely on 2D or partial 3D representations, limiting their ability to capture full scene geometry and constraining trajectory precision. We present GMT, a multimodal transformer framework that generates realistic and goal-directed object trajectories by jointly leveraging 3D bounding box geometry, point cloud context, semantic object categories, and target end poses. The model represents trajectories as continuous 6-DOF pose sequences and employs a tailored conditioning strategy that fuses geometric, semantic, contextual, and goaloriented information. Extensive experiments on synthetic and real-world benchmarks demonstrate that GMT outperforms state-of-the-art human motion and human-object interaction baselines, such as CHOIS and GIMO, achieving substantial gains in spatial accuracy and orientation control. Our method establishes a new benchmark for learningbased manipulation planning and shows strong generalization to diverse objects and cluttered 3D environments. Project page: https://huajian- zeng.github. io/projects/gmt/.

机器人操控6-DOF轨迹多模态3D场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。