arXiv:2604.09036cs.RO2026-04

用智能体自动生成可操作的机器人抓取场景,解决环境不现实导致任务失败的问题。

V-CAGE: Vision-Closed-Loop Agentic Generation Engine for Robotic Manipulation

  • 引入视觉闭环智能体,结合语义推理与物理交互生成合理场景
  • 通过视觉自检机制消除90%以上无效轨迹,提升数据可用性
  • 支持大规模数据压缩,存储减少90%仍保训练效果,适合机器人训练

扩展视觉-语言-动作(VLA)模型需要海量语义一致且物理可行的数据集。现有场景生成方法常缺乏上下文感知,难以合成高保真、富含语义信息的环境,导致目标位置不可达,任务提前失败。本文提出V-CAGE(视觉闭环智能体生成引擎),一个用于自主机器人数据合成的智能体框架。不同于传统脚本化流程,V-CAGE作为具身智能体系统,利用基础模型连接高层语义推理与底层物理交互。我们提出基于图像修复引导的场景构建方法,系统化布局上下文感知结构,确保生成场景语义有序且运动可达。为保障轨迹正确性,融合功能元数据与基于视觉-语言模型的闭环验证机制,充当视觉评判者,严格剔除隐蔽错误,切断错误传播链。最后,为突破大规模视频数据集的存储瓶颈,实现感知驱动压缩算法,文件大小缩减超90%,未影响下游VLA训练性能。通过集中语义布局规划与视觉自验证,V-CAGE实现了端到端自动化,支持多样、高质量机器人操作数据的大规模合成。

原文摘要 · Abstract (English)

Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing scene generation methods often lack context-awareness, making it difficult to synthesize high-fidelity environments embedded with rich semantic information, frequently resulting in unreachable target positions that cause tasks to fail prematurely. We present V-CAGE (Vision-Closed-loop Agentic Generation Engine), an agentic framework for autonomous robotic data synthesis. Unlike traditional scripted pipelines, V-CAGE operates as an embodied agentic system, leveraging foundation models to bridge high-level semantic reasoning with low-level physical interaction. Specifically, we introduce Inpainting-Guided Scene Construction to systematically arrange context-aware layouts, ensuring that the generated scenes are both semantically structured and kinematically reachable. To ensure trajectory correctness, we integrate functional metadata with a Vision-Language Model based closed-loop verification mechanism, acting as a visual critic to rigorously filter out silent failures and sever the error propagation chain. Finally, to overcome the storage bottleneck of massive video datasets, we implement a perceptually-driven compression algorithm that achieves over 90\% filesize reduction without compromising downstream VLA training efficacy. By centralizing semantic layout planning and visual self-verification, V-CAGE automates the end-to-end pipeline, enabling the highly scalable synthesis of diverse, high-quality robotic manipulation datasets.

机器人数据生成智能体视觉闭环

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。