arXiv:2608.29078cs.RO2026-08

无需人工示范,用语言指令自动生成机器人操作数据。

DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation

论文配图:DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation
图 1 · 摘自论文原文
  • 输入场景和语言指令,自动构建任务目标与成功标准。
  • 生成带随机物体配置的可行机械臂轨迹,验证后转为训练数据。
  • 实机测试表明效果优于直接部署,成本远低于人工采集。

视觉-语言-动作(VLA)模型在语言驱动的机器人操作中取得显著进展,但提升新工作空间中的性能仍常需该环境的动作标注数据。通过人工远程操控收集此类数据成本高昂,尤其当每个工作空间、物体布局或任务都需新示范时。我们提出DREAM框架,仅需捕获的工作空间图像和语言指令,即可生成用于微调预训练VLA的标注数据,无需特定任务的人工示范。DREAM首先重建工作空间,利用大语言模型将指令自动转化为符号化任务目标与成功标准,并通过任务与运动规划生成可行机器人轨迹。这些轨迹在随机化的物体配置下进行增强,由生成的成功标准验证后,渲染为图像-动作样本,用于VLA微调。通过真实机器人在语言驱动操作任务上的实验,我们评估DREAM能否作为部署环境的可扩展数据采集系统:比较其生成数据微调后的成功率是否优于直接部署,以及数据采集成本相较于人工操控的优劣。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have made strong progress in language-conditioned robot manipulation, but improving their performance in a new workspace still often requires action-labeled data from that environment. Collecting such data by human teleoperation is costly, especially when each workspace, object arrangement, or task may require new demonstrations. We present DREAM, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration. DREAM reconstructs the workspace, automatically translates the instruction into symbolic task goals and success criteria using a large language model, and uses task-and-motion planning to generate feasible robot trajectories. The planned trajectories are augmented across randomized object configurations, verified by the generated success criteria, and rendered into image-action examples for VLA fine-tuning. Through real-robot experiments on language-conditioned manipulation tasks, we study whether DREAM can serve as a scalable data-collection system for the deployment workspace by examining whether fine-tuning on its automatically generated data improves success over direct deployment and how its data-collection cost compares with human teleoperation when adapting a VLA to a new workspace.

机器人操作自动生成VLA模型零样本适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。