arXiv:2511.05855cs.RO2025-11AAAI被引 2

用视觉语言模型分解任务,让机器人学会精细操作。

Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic Skills

  • 通过视觉语言模型拆解复杂任务为原子技能,模拟训练基础策略。
  • 在仿真中训练带力约束的策略,避免操作时损坏物体。
  • 无需真人示范,可推广到多种新任务,适合工业级机器人应用。

自主执行长时程、高接触频率的操作任务通常需要大量真实世界数据和专家工程设计,带来高昂成本与扩展性挑战。本文提出一种新框架,融合层次化语义分解、强化学习(RL)、视觉语言模型(VLM)与知识蒸馏,以克服上述限制。将复杂任务分解为原子技能,每个基础动作由强化学习在仿真环境中独立训练,其训练过程显式引入力约束,防止敏感交互中造成物体损伤。视觉语言模型负责高层任务分解与技能规划,生成多样化专家示范。这些示范通过视觉-触觉扩散策略进行知识蒸馏,形成统一策略实现端到端执行。我们开展全面消融实验,评估不同VLM任务规划器对示范生成的影响,并系统比较多种模仿学习算法在技能蒸馏中的表现。大量仿真测试及物理部署验证表明,该方法可在不依赖昂贵人工示范的前提下完成长时程操作策略学习,且基于VLM引导的原子技能框架具备良好泛化能力,适用于多样任务场景。

原文摘要 · Abstract (English)

Autonomous execution of long-horizon, contact-rich manipulation tasks traditionally requires extensive real-world data and expert engineering, posing significant cost and scalability challenges. This paper proposes a novel framework integrating hierarchical semantic decomposition, reinforcement learning (RL), visual language models (VLMs), and knowledge distillation to overcome these limitations. Complex tasks are decomposed into atomic skills, with RL-trained policies for each primitive exclusively in simulation. Crucially, our RL formulation incorporates explicit force constraints to prevent object damage during delicate interactions. VLMs perform high-level task decomposition and skill planning, generating diverse expert demonstrations. These are distilled into a unified policy via Visual-Tactile Diffusion Policy for end-to-end execution. We conduct comprehensive ablation studies exploring different VLM-based task planners to identify optimal demonstration generation pipelines, and systematically compare imitation learning algorithms for skill distillation. Extensive simulation experiments and physical deployment validate that our approach achieves policy learning for long-horizon manipulation without costly human demonstrations, while the VLM-guided atomic skill framework enables scalable generalization to diverse tasks.

机器人操作视觉语言模型强化学习技能分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。