arXiv:2512.13090cs.RO2025-12中稿 · ICRA

用视觉语言模型和热启发扩散,实现多机器人语义导航与避障。

Multi-Robot Motion Planning from Vision and Language using Heat-Inspired Diffusion

  • 融合CLIP语义先验与碰撞规避扩散核,端到端生成语义指令轨迹。
  • 真实地图上成功率超基线,规划延迟降低30%以上。
  • 无需显式障碍物信息,适合复杂动态场景的机器人协同任务。

扩散模型近年来在机器人路径规划中展现出强大能力,能捕捉可行轨迹的多模态分布。然而,在灵活语言指令下的多机器人场景中,其应用仍受限。现有方法推理成本高,泛化性差,需显式构建环境表示且缺乏几何可达性推理机制。为此,我们提出语言条件热启发扩散(LHD),一种基于视觉的端到端框架,可生成语义约束、无碰撞的多机器人轨迹。LHD融合来自CLIP的语义先验与作为物理归纳偏置的碰撞规避扩散核,使规划器能在可到达工作空间内严格解析语言指令。该机制自然处理可达性方面的分布外(OOD)情况,引导机器人选择符合语义意图的可及替代路径,且推理时无需显式障碍物信息。在多种真实世界模拟地图上的广泛评估,以及真实机器人实验均表明,LHD在成功率上持续优于先前扩散规划方法,同时降低规划延迟。项目页面见:https://jebeom.github.io/lhd_project_page/

原文摘要 · Abstract (English)

Diffusion models have recently emerged as powerful tools for robot motion planning by capturing the multi-modal distribution of feasible trajectories. However, their extension to multi-robot settings with flexible, language-conditioned task specifications remains limited. Furthermore, current diffusion-based approaches incur high computational cost during inference and struggle with generalization because they require explicit construction of environment representations and lack mechanisms for reasoning about geometric reachability. To address these limitations, we present Language-conditioned Heat-inspired Diffusion (LHD), an end-to-end vision-based framework that generates language-conditioned, collision-free trajectories. LHD integrates semantic priors from CLIP, a vision-language model (VLM), with a collision-avoiding diffusion kernel serving as a physical inductive bias that enables the planner to interpret language commands strictly within the reachable workspace. This naturally handles out-of-distribution (OOD) scenarios -- in terms of reachability -- by guiding robots toward accessible alternatives that match the semantic intent, while eliminating the need for explicit obstacle information at inference time. Extensive evaluations on diverse real-world-inspired maps, along with real-robot experiments, show that LHD consistently outperforms prior diffusion-based planners in success rate, while reducing planning latency. Project page is available at: https://jebeom.github.io/lhd_project_page/

多机器人扩散模型视觉语言路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。