arXiv:2603.16028cs.RO2026-03被引 1

让大模型学会穿越多个狭窄开口的连续运动规划。

Geometry-Aligned LLM Fine-Tuning for Sequential Narrow-Opening Planning

  • 用失败反馈指导微调,生成可执行的路径点序列
  • 通过几何验证优化路径,成功率达92%以上
  • 适合需要长程空间推理的机器人任务

我们研究刚体通过多个连续狭窄开口的运动规划问题,该任务需要长时程几何推理,因为早期开口的配置会限制后续开口可达状态集。为此,提出一种几何对齐的大语言模型微调框架,生成固定长度、机器可读的路径点序列,确保跨开口的几何可行性与协调性。方法采用双层训练流程:首先基于人类示范进行故障驱动的LoRA监督微调(SFT),融入结构化失败反馈以学习常见错误模式并强制输出格式;其次使用分组相对策略优化(GRPO)进一步精炼相同LoRA适配器,通过模型预测规划器对路径点序列进行密化,并以确定性几何奖励评分,实现连续运动可行性。在仿真中定量与定性验证表明,本方法在分布内和分布外环境中均达到最高成功率,且能通过选择利于后续入口的出口姿态,体现长时程几何推理能力。

原文摘要 · Abstract (English)

We study rigid-body motion planning through multiple sequential narrow openings, which requires long-horizon geometric reasoning because the configuration used to traverse an early opening constrains the set of reachable configurations for subsequent ones. To achieve this, we propose a geometry-aligned large language model (LLM) fine-tuning framework that generates fixed-length, machine-readable waypoint sequences that are both geometrically feasible and coordinated across openings. Our approach uses a bi-level training pipeline. First, we perform failure-driven LoRA supervised fine-tuning (SFT) on human demonstrations, which incorporates structured failure feedback to teach the model common failure modes and enforce the output format. Second, we refine the same LoRA adapters using Group Relative Policy Optimization (GRPO) with geometric verification: each sampled waypoint sequence is densified by a model-based planner and scored with a deterministic geometry-derived reward to achieve continuous-motion feasibility. To validate the effectiveness of our proposed method, we provide both quantitative and qualitative results from simulations. Our method achieves the highest success rate in both in-distribution and out-of-distribution environments and qualitatively exhibits long-horizon geometric reasoning by selecting exit poses that facilitate entry into subsequent openings.

运动规划大模型几何推理机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。