让机器人听懂多指令并自动规划安全轨迹
GELATO: Multi-Instruction Trajectory Reshaping via Geometry-Aware Multiagent-based Orchestration
- 用视觉语言模型识别环境6维几何体,生成可验证约束
- 通过几何向量场优化实现平滑、安全的路径重规划
- 多智能体协同处理复杂指令,无需重新训练
我们提出GELATO——首个基于语言的轨迹重规划框架,融合几何环境感知与多智能体反馈协调,支持人机交互中的多指令场景。不同于以往学习方法,本方案通过VLM辅助的多视角管道自动将场景物体注册为6D几何原型,并由大语言模型将自由形式的多指令转化为明确、可验证的几何约束。这些约束被整合进几何感知的向量场优化中,以在保持路径平滑性、可行性与避障距离的前提下调整初始轨迹。我们进一步引入基于观察者的多智能体协调机制,有效处理多指令输入及目标间交互,显著提升成功率且无需重新训练。仿真与真实世界实验表明,该方法在轨迹平滑性、安全性与可解释性方面均优于现有最优基线。
原文摘要 · Abstract (English)
We present GELATO -- the first language-driven trajectory reshaping framework to embed geometric environment awareness and multi-agent feedback orchestration to support multi-instruction in human-robot interaction scenarios. Unlike prior learning-based methods, our approach automatically registers scene objects as 6D geometric primitives via a VLM-assisted multi-view pipeline, and an LLM translates free-form multiple instructions into explicit, verifiable geometric constraints. These are integrated into a geometric-aware vector field optimization to adapt initial trajectories while preserving smoothness, feasibility, and clearance. We further introduce a multi-agent orchestration with observer-based refinement to handle multi-instruction inputs and interactions among objectives -- increasing success rate without retraining. Simulation and real-world experiments demonstrate our method achieves smoother, safer, and more interpretable trajectory modifications compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。