arXiv:2607.08024cs.CVcs.AI2026-07

让机器人在复杂环境中自主规划,兼顾语义理解与空间可行性。

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

论文配图:APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
图 1 · 摘自论文原文
  • 用语言和视觉交替思考,动态调整推理方式。
  • 在厨房场景中成功率显著提升,尤其在空间受限时表现更优。
  • 适合需要长期规划的机器人任务,如家庭服务机器人。

长时序机器人规划需同时考虑任务语义结构与几何可行性。为成功执行任务,机器人必须分解目标、选择相关物体并排序动作,同时确保计划满足空间约束,如可用空间有限和物体碰撞避免。本文提出APIVOT,一种基于视觉-语言模型(VLM)的规划器,通过自适应地交错使用语言与视觉思考来实现长时序规划。APIVOT利用语言进行语义推理,同时将视觉思考作为未来状态的内在模拟,以验证几何可行性。在长时序厨房任务上,APIVOT优于通用VLM和现有规划框架,在空间受限条件下取得最大提升。实验表明,APIVOT学会了有意义的模态选择行为,证明视觉-语言思考的自适应交错能有效提升规划成功率与推理效率。

原文摘要 · Abstract (English)

Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.

机器人规划视觉语言模型自适应推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。