打造四足机器人通用运动控制框架,实现真实场景下的智能行走与交互。
Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

- 构建多源数据融合管道,生成1.6万段物理可行的动作片段
- 首次发现四足运动追踪的规模效应,训练越强零样本追踪能力越强
- 集成地形预测与激光记忆,实现复杂地形稳健行进
运动控制器是具身智能系统中最基础的模块。尽管基于大规模人体动作捕捉数据和运动追踪范式的类人机器人控制近年取得显著进展,但将此方法迁移到四足场景仍面临挑战:动物动作数据稀缺且难规模化采集,跨形态重定向仍不稳定。本文提出ABot-C0,一种面向四足机器人的通用运动控制体系,建立三大行为基础:可扩展的多源动作数据流水线、跨运动追踪、移动与环境交互的鲁棒策略学习,以及统一部署栈以保障真实世界稳定运行。核心上,通过条件视频生成合成、标注动作捕捉、远程操控与人类设计构建数据金字塔,生成16,074段物理可行的动作片段,作为多样化运动学习的基础。基于大规模动作数据,采用流匹配通用策略首次揭示四足运动追踪的规模定律:性能随训练规模提升而持续增强,并具备零样本追踪未见动作的能力。进一步通过三阶段特权到感知框架,结合时间序列激光雷达记忆与地形预测监督,实现稳健全地形移动。上述组件共同构成一个运动通用智能体,协调多策略执行、平滑行为切换、节能控制与安全机制,适用于真实部署。在城市地形自主导航与陪伴式多模态交互的大量实验中,验证了四足机器人正从功能演示迈向产品级行为智能。
原文摘要 · Abstract (English)
The motion controller is one of the most fundamental modules in embodied intelligence systems. Driven by large-scale human motion-capture data and the motion-tracking paradigm, humanoid control has achieved remarkable progress in recent years. However, migrating this recipe to the quadrupedal setting is far less straightforward: animal motion data is scarcer and harder to capture at scale than human data, and cross-embodiment retargeting remains fragile. We present ABot-C0, a generalist motion-control system for quadruped robots that establishes three complementary behavior foundations: a scalable multi-source motion-data pipeline, robust policy learning across motion tracking, locomotion, and scene interaction, and a unified deployment stack for reliable real-world operation. Fundamentally, we construct a data pyramid through conditional video-generation synthesis, annotated motion capture, teleoperation, and human design, producing 16,074 physically feasible motion clips as the data foundation for diverse motion-learning demands. With large-scale motion data, a Flow-Matching generalist policy demonstrates, for the first time, a scaling law for quadruped motion tracking: performance improves consistently as training scales up, with zero-shot capability to track unseen motions. We then go a step further toward robust all-terrain locomotion by adopting a three-stage privileged-to-perceptive framework with temporal LiDAR memory and terrain-predictive supervision. Collectively, these components form a motion generalist that coordinates multi-policy execution, smooth behavior transitions, energy-efficient control, and safety mechanisms for real-world deployment. Extensive experiments on urban-terrain autonomous navigation and companion-style multimodal interaction demonstrate that quadruped robots can move beyond functional demos toward product-level behavioral intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。