用多尺度意图扩散让机器人听懂文字指令并自然行动
MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

- 通过分解行为意图,分层次引导动作生成
- 在多个数据集上实现更自然、更符合语义的物理仿真动作
- 适合研究人形机器人控制与文本驱动生成的学者
让基于物理的人形机器人根据高层文本指令执行多样化行为仍是重大挑战。现有方法要么采用先生成运动再跟踪的两阶段范式,存在运动与物理控制之间的领域偏移;要么采用端到端模仿学习直接从文本生成动作,但面临文本与低层动作间的显著模态差距,导致语义对齐困难。值得注意的是,人形机器人状态蕴含丰富的运动动态,比低层动作更接近文本描述的语义,可作为语义桥梁。为此,我们提出MIND,一种新颖的端到端扩散框架,利用行为意图作为文本与低层动作之间的语义纽带。核心是多尺度意图扩散机制:全局意图预测器捕捉整体行为动态以指导整体合成,即时意图预测器在每一步提供细粒度信号以优化局部行为。这种分层意图设计为机器人控制引入结构化归纳偏置,增强语义对齐与行为自然性。此外,MIND将人形状态编码到潜在空间,提升意图建模效果。大量实验表明,MIND优于现有方法,在多个数据集上均能从文本命令生成连贯、物理合理且语义一致的行为。
原文摘要 · Abstract (English)
Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typically follow either a two-stage paradigm that combines kinematic motion generation with physics-based tracking, or an end-to-end imitation-learning paradigm that directly generates actions from text. However, the former suffers from the inherent domain shift between kinematic generation and physics-based tracking, while the latter struggles with the substantial modality gap between textual commands and low-level actions, limiting effective semantic alignment. Notably, humanoid states encode rich motion dynamics that are more semantically aligned with textual descriptions than low-level actions, making them a natural basis for deriving behavioral intent. Building upon this insight, we propose MIND, a novel end-to-end diffusion framework for text-driven physics-based humanoid control that leverages behavioral intent as a semantic bridge between textual commands and low-level actions. At its core, MIND introduces a multi-scale intent diffusion mechanism, where a holistic intent predictor captures global behavioral dynamics to guide overall behavior synthesis, while an immediate intent predictor provides step-wise, fine-grained signals for local behavior refinement at each diffusion step. This hierarchical intent formulation imposes a structured inductive bias for humanoid control, improving semantic alignment and behavioral naturalness. Furthermore, MIND encodes humanoid states into a latent space to enable more effective semantic intent modeling. Extensive experiments demonstrate that MIND outperforms existing methods and synthesizes coherent, physically plausible, and semantically aligned humanoid behaviors from text commands. Project page: https://binlee26.github.io/MIND_page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。