机器人自主生成交互数据,无需人工干预,可自动重置环境并持续学习。
RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset
- 用视觉语言模型规划任务,图神经网络执行动作,实现闭环自主采集
- 模拟中复杂任务成功率高达90%,远超传统方法的接近零表现
- 适合需要大规模真实交互数据的机器人研发团队快速部署
获取大规模物理交互数据是现代机器人学习的关键前提,但传统人工参与的数据采集方式成本高、难扩展。为此,我们提出完全自主的闭环数据生成系统RADAR,彻底摆脱人工干预。RADAR将认知任务分解为四模块:基于2-5个3D人类示范作为几何先验,视觉语言模型通过精准语义对象定位与技能检索生成场景相关任务;图神经网络策略利用上下文模仿学习将子任务转化为物理动作;执行后,视觉语言模型通过结构化视觉问答流程自动评估成功与否;最后,有限状态机协同实现自主环境重置与非对称数据路由。系统采用前向-反向联合规划,严格遵循后进先出因果序列,能无缝恢复杂乱工作区并有效应对执行失败。这种持续脑-小脑协同机制使数据采集成为自维持过程。大量实验表明,该框架在仿真中对复杂长程任务成功率可达90%,显著优于传统基线;在真实场景中,仅需少量样本即可可靠执行多样接触型技能(如柔性物体操作),无需领域特定微调,提供高度可扩展的机器人数据采集范式。
原文摘要 · Abstract (English)
The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms. To break this barrier, we introduce Robust Autonomous Data Acquisition for Robotics (RADAR), a fully autonomous, closed-loop data generation engine that completely removes human intervention from the collection cycle. RADAR elegantly divides the cognitive load into a four-module pipeline. Anchored by 2-5 3D human demonstrations as geometric priors, a Vision-Language Model first orchestrates scene-relevant task generation via precise semantic object grounding and skill retrieval. Next, a Graph Neural Network policy translates these subtasks into physical actions via in-context imitation learning. Following execution, the VLM performs automated success evaluation using a structured Visual Question Answering pipeline. Finally, to shatter the bottleneck of manual resets, a Finite State Machine orchestrates an autonomous environment reset and asymmetric data routing mechanism. Driven by simultaneous forward-reverse planning with a strict Last-In, First-Out causal sequence, the system seamlessly restores unstructured workspaces and robustly recovers from execution failures. This continuous brain-cerebellum synergy transforms data collection into a self-sustaining process. Extensive evaluations highlight RADAR's exceptional versatility. In simulation, our framework achieves up to 90% success rates on complex, long-horizon tasks, effortlessly solving challenges where traditional baselines plummet to near-zero performance. In real-world deployments, the system reliably executes diverse, contact-rich skills (e.g., deformable object manipulation) via few-shot adaptation without domain-specific fine-tuning, providing a highly scalable paradigm for robotic data acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。