arXiv:2510.14968cs.ROcs.AI2025-10NeurIPS被引 4

用检索方法自动分解任务演示,提升长程操作的规划准确性。

RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks

  • 通过视觉特征对齐,从训练数据中检索匹配的子任务片段。
  • 在仿真与真实场景中均超越当前最佳方法,性能更稳定。
  • 适合需要精准任务分解的机器人长程操作研究者。

为应对长程任务挑战,现有分层视觉-语言-动作(VLAs)框架利用基于视觉语言模型(VLM)的规划器,将复杂操作任务分解为低层视觉运动策略可处理的简单子任务。通常,该规划器需微调以学习任务分解,依赖人工标注或启发式规则对目标任务演示进行分段。然而,启发式生成的子任务可能与低层策略训练数据差异较大,导致性能下降。为此,我们提出基于检索的演示分解器(RDD),通过将分解后的子任务区间视觉特征与底层视觉运动策略训练数据中的特征对齐,实现自动分解。实验表明,该方法在仿真和真实世界任务中均优于现有最优子任务分解方法,展现出跨多种场景的鲁棒性。代码与更多结果见 rdd-neurips.github.io。

原文摘要 · Abstract (English)

To tackle long-horizon tasks, recent hierarchical vision-language-action (VLAs) frameworks employ vision-language model (VLM)-based planners to decompose complex manipulation tasks into simpler sub-tasks that low-level visuomotor policies can easily handle. Typically, the VLM planner is finetuned to learn to decompose a target task. This finetuning requires target task demonstrations segmented into sub-tasks by either human annotation or heuristic rules. However, the heuristic subtasks can deviate significantly from the training data of the visuomotor policy, which degrades task performance. To address these issues, we propose a Retrieval-based Demonstration Decomposer (RDD) that automatically decomposes demonstrations into sub-tasks by aligning the visual features of the decomposed sub-task intervals with those from the training data of the low-level visuomotor policies. Our method outperforms the state-of-the-art sub-task decomposer on both simulation and real-world tasks, demonstrating robustness across diverse settings. Code and more results are available at rdd-neurips.github.io.

任务分解机器人规划视觉语言模型长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。