arXiv:2607.04880cs.RO2026-07

仅用一张图和指令就能生成个性化机器人训练数据。

PRISM: Personalized Robotic Dataset Generation via Image-based Scene and Motion Synthesis

论文配图:PRISM: Personalized Robotic Dataset Generation via Image-based Scene and Motion Synthesis
图 1 · 摘自论文原文
  • 输入单张图像+自然语言指令,自动生成匹配环境的虚拟场景
  • 在LIBERO任务中达100%成功率,且泛化能力更强
  • 无需人工操作,适合家庭、工厂等个性化场景应用

大型预训练视觉-语言-动作模型虽提升了机器人策略学习效果,但直接部署于用户特定环境仍面临泛化不足问题,需收集适配目标环境的数据。远程操控可获得对齐数据,但成本高且难扩展;仿真易扩展,却难以贴近真实环境并生成任务相关轨迹。为此,我们提出PRISM,一个端到端管道,仅需一张图像和自然语言指令即可生成个性化机器人数据集。PRISM构建与用户环境语义和几何对齐的数字孪生场景,同时在实例层面保持多样性,并合成无需人工干预的可执行示范。大量实验表明,基于PRISM生成数据训练的策略在LIBERO和LIBERO-Plus上优于基线,三类真实操作任务实现100%成功率,且在与训练环境不同的场景中仍保持强性能。

原文摘要 · Abstract (English)

Recent advances in large-scale pretrained vision-language-action models have improved robot policy learning, but directly deploying such policies in user-specific environments remains challenging due to limited generalization, which inevitably requires collecting a dataset tailored to the target environment. Teleoperation yields well-aligned data but is costly and difficult to scale, whereas simulation scales easily but struggles to resemble the target environment and generate task-specific trajectories. To meet both simultaneously, we propose PRISM, an end-to-end pipeline that generates personalized robotic datasets from a single image and a natural-language instruction. PRISM constructs digital cousin scenes that are semantically and geometrically aligned with the user environment yet diverse at the instance level, and synthesizes executable demonstrations without human teleoperation. Extensive experiments show that policies trained on PRISM-generated datasets outperform those trained on baseline-generated datasets on LIBERO and LIBERO-Plus, achieve up to 100\% success rate on three real-world manipulation tasks, and maintain stronger performance when evaluated in environments that differ from those seen during training.

机器人数据生成个性化虚拟仿真视觉-语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。