让机器人听懂变化的指令,实时调整动作策略。
Vision-Language-Policy Model for Dynamic Robot Task Planning
- 基于视觉-语言模型,融合多模态感知与任务推理生成行为策略。
- 可在任务执行中动态响应指令变化,实现策略自适应更新。
- 跨机器人平台通用,在真实场景中展现强泛化能力。
在非结构化环境中,如何将自然语言指令与自主执行有效衔接仍是机器人领域的挑战。这要求机器人能通过多模态感知理解当前任务场景,并规划行为以达成目标。传统任务规划方法常难以连接低层执行与高层推理,且在执行过程中无法动态调整策略以应对指令变更,限制了其灵活性与适应性。本文提出一种基于语言模型的动态机器人任务规划新框架——视觉-语言-策略(Vision-Language-Policy, VLP)模型。该模型基于在真实数据上微调的视觉-语言模型,可解析语义指令,结合对当前任务场景的推理,生成控制机器人完成任务的行为策略。此外,它能根据任务变化动态调整策略,实现灵活适应。在多种机器人和真实任务上的实验表明,该模型能高效适应新场景并动态更新策略,展现出强大的规划自主性与跨平台泛化能力。
原文摘要 · Abstract (English)
Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene through multiple modalities, and to plan their behaviors to achieve their intended goals. Traditional robotic task-planning approaches often struggle to bridge low-level execution with high-level task reasoning, and cannot dynamically update task strategies when instructions change during execution, which ultimately limits their versatility and adaptability to new tasks. In this work, we propose a novel language model-based framework for dynamic robot task planning. Our Vision-Language-Policy (VLP) model, based on a vision-language model fine-tuned on real-world data, can interpret semantic instructions and integrate reasoning over the current task scene to generate behavior policies that control the robot to accomplish the task. Moreover, it can dynamically adjust the task strategy in response to changes in the task, enabling flexible adaptation to evolving task requirements. Experiments conducted with different robots and a variety of real-world tasks show that the trained model can efficiently adapt to novel scenarios and dynamically update its policy, demonstrating strong planning autonomy and cross-embodiment generalization. Videos: https://robovlp.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。