arXiv:2602.12215cs.RO2026-02中稿 · RSS 2026, Project …被引 39

用统一数据流训练10亿参数机器人模型,能高效利用低质量轨迹提升性能。

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

  • 通过结构化潜空间学习动态,联合建模动作与视觉预测,区分不同数据质量。
  • 在30小时真实/模拟轨迹上训练,接触、灵巧、长时任务分别提升21%~48%。
  • 支持数据高效微调,仅用30%低质数据就增益10%,适合资源有限场景。

近期机器人基础模型主要依赖大规模行为克隆,虽能模仿专家动作,却丢弃了异构具身数据中蕴含的可迁移动态知识。尽管统一世界模型(UWM)具备利用此类多样化数据的潜力,但现有实现因数据使用粗略、数据集碎片化,难以扩展至基础模型规模。本文提出LDA-1B,一种通过通用具身数据摄入实现规模化的新模型,联合学习动态、策略与视觉预测,并为不同质量的数据分配特定角色。为支撑该范式,我们构建并标准化了包含超过30,000小时人类与机器人轨迹的EI-30k数据集,统一格式。在结构化DINO潜空间中进行预测,避免冗余像素级外观建模,实现对异构数据的可扩展动态学习。同时,LDA-1B采用多模态扩散变压器处理异步视觉与动作流,支持10亿参数量级稳定训练。仿真与真实世界实验表明,该模型在接触密集、灵巧操作和长时程任务上分别优于先前方法(如$π_{0.5}$)达21%、48%和23%。值得注意的是,仅利用30%通常被丢弃的低质量轨迹,模型即可实现10%的性能增益。

原文摘要 · Abstract (English)

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing instantiations struggle to scale to foundation-level due to coarse data usage and fragmented datasets. We introduce LDA-1B, a robot foundation model that scales through universal embodied data ingestion by jointly learning dynamics, policy, and visual forecasting, assigning distinct roles to data of varying quality. To support this regime at scale, we assemble and standardize EI-30k, an embodied interaction dataset comprising over 30k hours of human and robot trajectories in a unified format. Scalable dynamics learning over such heterogeneous data is enabled by prediction in a structured DINO latent space, which avoids redundant pixel-space appearance modeling. Complementing this representation, LDA-1B employs a multi-modal diffusion transformer to handle asynchronous vision and action streams, enabling stable training at the 1B-parameter scale. Experiments in simulation and the real world show LDA-1B outperforms prior methods (e.g., $π_{0.5}$) by up to 21\%, 48\%, and 23\% on contact-rich, dexterous, and long-horizon tasks, respectively. Notably, LDA-1B enables data-efficient fine-tuning, gaining 10\% by leveraging 30\% low-quality trajectories typically harmful and discarded.

机器人具身智能扩散模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。