让机器人提前理解动作意图,用语言指挥更自然的全身动作。
Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

- 分层架构预测未来动作意图,连接语言与实时控制
- 在HumanML3D上达94.42%成功率,BABEL数据集子序列FID低至0.152
- 适合需要流畅、预判式人形机器人控制的研究与应用
自然语言是人形机器人直观的控制接口,但实现全身流式控制需具备当前可执行性与未来物理变化的预判能力。现有语言驱动系统通常生成需底层追踪器被动修正的运动参考,或使用不显式编码未来接触变化、转移和平衡准备的隐状态/动作策略。本文提出DAJI(动态对齐关节意图)框架,通过层次化设计建立语言生成与闭环控制间的预判性关节意图接口。DAJI-Act通过学生驱动的轨迹回放,将未来感知的教师模型压缩为可部署的扩散动作策略;DAJI-Flow则基于语言和意图历史自回归生成未来意图片段。实验表明,DAJI在预判性隐空间学习、单指令生成与流式指令跟随任务中表现优异,在HumanML3D风格生成任务中达到94.42%的回放成功率,在BABEL数据集上实现0.152的子序列FID。
原文摘要 · Abstract (English)
Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable now and anticipatory of future physical transitions. Existing language-conditioned humanoid systems typically generate kinematic references that a low-level tracker must repair reactively, or use latent/action policies whose outputs do not explicitly encode upcoming contact changes, support transfers, and balance preparation. We propose \textbf{DAJI} (\emph{Dynamics-Aligned Joint Intent}), a hierarchical framework that learns an anticipatory joint-intent interface between language generation and closed-loop control. DAJI-Act distills a future-aware teacher into a deployable diffusion action policy through student-driven rollouts, while DAJI-Flow autoregressively generates future intent chunks from language and intent history. Experiments show that DAJI achieves strong results in anticipatory latent learning, single-instruction generation, and streaming instruction following, reaching 94.42\% rollout success on HumanML3D-style generation and 0.152 subsequence FID on BABEL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。