ABot-M0通过动作流形学习,让机器人在多形态平台上高效执行任务。
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
- 构建统一数据集与模型架构,实现跨平台数据融合
- 提出动作流形假设,提升动作预测速度与稳定性
- 支持模块化视觉感知,增强三维空间理解能力
构建通用具身智能体面临硬件多样性的挑战,常被概括为“一脑多形”范式。进展受限于数据碎片化、表示不一致及训练目标错位。本文提出ABot-M0框架,通过系统性数据清洗与标准化,整合六大数据集,构建包含超过600万轨迹、9500小时数据的UniACT-dataset,覆盖多种机器人形态与任务场景。统一预训练显著提升跨平台与任务的知识迁移与泛化能力。为提升动作预测效率与稳定性,提出动作流形假设:有效动作位于由物理规律和任务约束支配的低维光滑流形上。基于此,设计动作流形学习(AML),采用DiT骨干网络直接预测连续动作序列,将学习从去噪转为投影至可行流形,提升解码速度与策略稳定性。通过双流机制实现模块化感知,融合视觉语言模型语义与几何先验,结合可插拔3D模块如VGGT和Qwen-Image-Edit,增强空间理解且不修改主干网络,缓解标准VLM在3D推理中的局限。实验表明各组件独立运作并具有叠加增益。代码与流程将全部开源,供后续研究复现。
原文摘要 · Abstract (English)
Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, many-forms'' paradigm. Progress is hindered by fragmented data, inconsistent representations, and misaligned training objectives. We present ABot-M0, a framework that builds a systematic data curation pipeline while jointly optimizing model architecture and training strategies, enabling end-to-end transformation of heterogeneous raw data into unified, efficient representations. From six public datasets, we clean, standardize, and balance samples to construct UniACT-dataset, a large-scale dataset with over 6 million trajectories and 9,500 hours of data, covering diverse robot morphologies and task scenarios. Unified pre-training improves knowledge transfer and generalization across platforms and tasks, supporting general-purpose embodied intelligence. To improve action prediction efficiency and stability, we propose the Action Manifold Hypothesis: effective robot actions lie not in the full high-dimensional space but on a low-dimensional, smooth manifold governed by physical laws and task constraints. Based on this, we introduce Action Manifold Learning (AML), which uses a DiT backbone to predict clean, continuous action sequences directly. This shifts learning from denoising to projection onto feasible manifolds, improving decoding speed and policy stability. ABot-M0 supports modular perception via a dual-stream mechanism that integrates VLM semantics with geometric priors and multi-view inputs from plug-and-play 3D modules such as VGGT and Qwen-Image-Edit, enhancing spatial understanding without modifying the backbone and mitigating standard VLM limitations in 3D reasoning. Experiments show components operate independently with additive benefits. We will release all code and pipelines for reproducibility and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。