arXiv:2603.18811cs.RO2026-03被引 1

用自然语言自动生成机器人仿真环境和操作轨迹,无需人工干预。

V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors

  • 结合大语言模型与3D生成模型构建物理合理的场景
  • 利用视频生成模型作为运动先验,生成可执行的机器人轨迹
  • 在真实机械臂上成功实现模拟到现实的迁移

训练通用机器人需要大规模、多样化的操作数据,但真实世界采集成本过高,现有模拟器又受限于固定资产库和手动规则。为此,我们提出V-Dreamer,一个完全自动化的框架,可直接从自然语言指令生成开放词汇、可模拟的操作环境和可执行的专家轨迹。V-Dreamer采用新型生成流程,利用大语言模型和3D生成模型构建物理可信的三维场景,并通过几何约束验证确保布局稳定无碰撞。关键在于行为合成阶段,我们借助视频生成模型作为丰富的运动先验,再通过基于CoTracker3和VGGT的稳健Sim-to-Gen视觉-运动学对齐模块,将视觉预测转化为可执行机器人轨迹。该流程无需人工干预即可实现高视觉多样性与物理保真度。为评估生成数据,我们在合成轨迹上训练模仿学习策略,涵盖多种物体与环境变化。在使用Piper机械臂的桌面操作任务中,实验表明策略能有效泛化至未见物体,在仿真中表现稳健,并成功实现模拟到现实的迁移,实际操控了新实物。

原文摘要 · Abstract (English)

Training generalist robots demands large-scale, diverse manipulation data, yet real-world collection is prohibitively expensive, and existing simulators are often constrained by fixed asset libraries and manual heuristics. To bridge this gap, we present V-Dreamer, a fully automated framework that generates open-vocabulary, simulation-ready manipulation environments and executable expert trajectories directly from natural language instructions. V-Dreamer employs a novel generative pipeline that constructs physically grounded 3D scenes using large language models and 3D generative models, validated by geometric constraints to ensure stable, collision-free layouts. Crucially, for behavior synthesis, we leverage video generation models as rich motion priors. These visual predictions are then mapped into executable robot trajectories via a robust Sim-to-Gen visual-kinematic alignment module utilizing CoTracker3 and VGGT. This pipeline supports high visual diversity and physical fidelity without manual intervention. To evaluate the generated data, we train imitation learning policies on synthesized trajectories encompassing diverse object and environment variations. Extensive evaluations on tabletop manipulation tasks using the Piper robotic arm demonstrate that our policies robustly generalize to unseen objects in simulation and achieve effective sim-to-real transfer, successfully manipulating novel real-world objects.

机器人视频生成仿真自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。