arXiv:2605.13775cs.ROcs.CV2026-05被引 1

用自进化框架让机器人在极少数据下学会复杂操作

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

论文配图:RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
图 1 · 摘自论文原文
  • 将视觉语言模型与视频生成器联动,形成自我优化循环
  • 仅用500张无标注图,性能超越监督学习基线50倍
  • 适合数据稀缺场景的机器人操控训练,抗遗忘能力强

机器人的可扩展操作受限于任务对齐的物理交互数据稀缺。尽管视觉-语言模型(VLM)和视频生成模型(VGM)有望实现自主数据合成,但分别存在语义-空间错位和物理幻觉问题。为此,我们提出RoboEvolve,将VLM规划器与VGM模拟器耦合,构建相互促进的共演化循环。系统基于未标注种子图像,采用受认知启发的双阶段机制:(i) 白天探索通过语义控制的多粒度奖励,发现物理合理的动作行为;(ii) 夜间巩固从“近失败”中挖掘经验以稳定策略优化。在自主渐进式课程引导下,系统可自然从基础原子动作扩展至复杂任务。大量实验表明,RoboEvolve (I) 效果更优,使基础规划器提升30个百分点,模拟器成功率平均提高48%;(II) 数据效率极高,仅需500张未标注种子即超越全监督基线,数据量减少50倍;(III) 展现出强健的持续学习能力,无灾难性遗忘。

原文摘要 · Abstract (English)

The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video generation models (VGMs) hold promise for autonomous data synthesis, they suffer from semantic-spatial misalignment and physical hallucinations, respectively. To bridge this gap, we introduce RoboEvolve, a novel framework that couples a VLM planner and a VGM simulator into a mutually reinforcing co-evolutionary loop. Operating purely on unlabeled seed images, RoboEvolve leverages a cognitive-inspired dual-phase mechanism: (i) daytime exploration fosters physically grounded behavioral discovery through a semantic-controlled multi-granular reward, and (ii) nighttime consolidation mines "near-miss" failures to stabilize policy optimization. Guided by an autonomous progressive curriculum, the system naturally scales from simple atomic actions to complex tasks. Extensive experiments demonstrate that RoboEvolve (I) achieves superior effectiveness, elevating base planners by 30 absolute points and amplifying simulator success by 48% on average; (II) exhibits extreme data efficiency, surpassing fully supervised baselines with merely 500 unlabeled seeds--a 50x reduction; and (III) demonstrates robust continual learning without catastrophic forgetting.

机器人操控自进化小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。