arXiv:2607.20345cs.ROcs.AI2026-07

用少量数据让机器人在超市真实场景中稳定完成补货任务

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

论文配图:Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
图 1 · 摘自论文原文
  • 通过数据高效微调与视觉聚焦,减少对大模型的依赖
  • 仅用一台显卡,实现从实验室失败到真实场景成功的转变
  • 适合关注机器人落地应用的开发者和研究者

视觉-语言-动作(VLA)人形机器人在从实验室到实际商店的部署中,仍面临执行错误、分布偏移和环境变化等挑战。本文提出DEED(数据高效后训练与经验驱动学习)框架,在使用Unitree G1-Edu人形机器人和GR00T N1.6基础模型的超市芯片补货任务中进行评估。该框架包含三部分:(1) 数据高效后训练流程,含控制频率对齐、数据筛选、任务相关视觉突出和降低VLA依赖;(2) 基于文本优势前缀和视觉-语言价值函数的经验驱动优化,源自RECAP;(3) 用于分析分布内/外行为的潜在空间分析工具。结果表明,缩小实验室到现实差距主要在于系统集成而非架构设计:精心设计的数据与针对性后训练,仅需单张GPU即可将原本微调失败的策略转化为可靠的真实世界系统。

原文摘要 · Abstract (English)

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.

人形机器人视觉语言动作落地应用数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。