SABER用真实超市数据提升机器人在零售场景的动手能力。
SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation

- 用头戴相机+360度摄像头采集真实超市中的人类操作行为。
- 数据包含44.8万样本,使机器人任务成功率超基线2.19倍。
- 适合做零售机器人、具身智能和动作迁移的研究者使用。
真实世界机器人部署依赖丰富的领域特定动作数据,而通用机器人基础模型在未见过的复杂任务(如零售环境操作)中表现不佳。根本原因在于数据缺口:零售环境未被通用预训练数据覆盖,且通过远程操控收集数据成本高、难扩展。我们提出SABER,一个基于多个真实生鲜超市中超过100小时自然录制的高保真零售机器人动作数据集。头戴式摄像机记录交互点处精细的手部动作,而DreamVu的ALIA相机提供全场景360度外视角,同步捕捉所有参与者与活动。该组合实现了人类零售行为的完整记录:灵巧手部动作、全身运动及场景动态,无需布景、脚本或远程操控。SABER包含44.8万个训练样本,涵盖三种动作表示流:25,000个通过LAPA风格编码的潜在动作序列,18,600个重映射至机器人关节空间的灵巧手姿轨迹,以及1,200个同步重映射至人形机器人的全身运动序列。在GR00T N1.6上采用共享主干多任务微调方案,SABER在十个零售操作任务上的平均成功率达29.3%,超过微调基线(13.4%)2.19倍。结果表明,更优的机器人能力可通过今日即可规模化采集的真实数据实现,无需机器人参与。数据与代码已开放:https://dreamvu.ai/saber
原文摘要 · Abstract (English)
Robotic deployment in real-world environments depends on rich, domain-specific action data as much as on strong model architecture. General-purpose robot foundation models show modest performance in complex unseen tasks such as manipulation in a retail domain when applied out of the box. The root cause is a data gap: retail environments are structurally absent from general robot pretraining distributions, and the path to filling that gap through teleoperation is prohibitively expensive, logistically constrained, and difficult to scale. We introduce SABER, a high-fidelity retail robotics action dataset built from over 100 hours of natural in-store capture across multiple real grocery environments. Egocentric footage from head-mounted cameras records fine-grained hand activity at the point of interaction, while exocentric 360-degree scene footage from DreamVu's ALIA camera simultaneously observes all actors and activities across the entire space. This combination yields a uniquely complete picture of human retail behavior: dexterous hand activity, whole-body motion, and scene dynamics, all captured without staging, scripting, or teleoperation overhead. The SABER corpus contains 44.8K training samples across three action representation streams: 25K latent action sequences via LAPA-style encoding, 18.6K dexterous hand-pose trajectories retargeted to robot joint space, and 1.2K whole-body synchronized motion sequences retargeted to a humanoid embodiment. When applied to GR00T N1.6 via a shared-backbone multi-task post-training recipe, SABER yields a mean success rate of 29.3% across ten retail manipulation tasks -- more than 2.19x over fine-tuning baselines (13.4%). SABER demonstrates that the path to capable retail robots runs through better data, which can be collected today, at scale, without a robot in the loop. The dataset and code are available at https://dreamvu.ai/saber
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。