从一张图生成可执行任务,自动优化失败案例。
Scene2Demo: Self-Evolving Embodied Data Generation via Object-Action Graph
- 用物体-动作图结构化任务,自动生成仿真场景和视频。
- 102个任务中执行成功率达71.6%,长任务成功率提升显著。
- 适合机器人数据生成与行为克隆研究者使用。
我们提出Scene2Demo,一种离线具身数据自演化生成框架。给定一张真实世界RGB图像和用户查询,该框架构建交互式仿真场景,生成可执行的任务配置、多视角执行视频及离线机器人学习数据集。通过物体-动作图的结构化多模块流程,以物体为中心配置和动作转移来表示任务生成。失败或不完整的执行由反馈代理通过视觉回放检查,并通过序列修改或参数调整优化动作流。在102个自动生成的初级场景-任务对上,Scene2Demo实现71.6%的执行成功率;在四个代表性长时任务中,自演化显著提升任务成功率和子任务执行质量。与RoboGen和GenSim2对比显示,在自动化数据生成设置下具备更强的任务规划与执行性能。行为克隆策略在两个代表性任务上分别达到96.0%和92.0%的成功率,验证了生成数据支持下游策略学习的有效性。
原文摘要 · Abstract (English)
We present Scene2Demo, a self-evolving framework for offline embodied data generation. Given a single real-world RGB image and a user query, Scene2Demo constructs an interactive simulated scene and generates executable task configurations, multi-view execution videos, and offline robot-learning datasets. Scene2Demo uses a structured multi-module workflow via an object-action graph, representing task generation through object-centric configurations and action transitions. Failed or incomplete executions are further refined by feedback agents that inspect visual rollouts and revise action flows through sequence modification or parameter adjustment. Across 102 automatically generated primitive scene-task pairs, Scene2Demo achieves a 71.6\% execution success rate; on four representative long-horizon tasks, self-evolution improves both task success and subtask-level execution quality over primitive-only execution, and comparisons with RoboGen and GenSim2 show stronger task planning and execution performance under automated data-generation settings. Finally, behavior cloning policies achieve 96.0\% and 92.0\% success on two representative tasks, validating that the generated data can support downstream policy learning. Our project page is available at https://scene2demo-anon.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。