用扩散模型生成多人类机器人物理协调动作,仅凭文本指令即可实现
InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
- 采用多流扩散变换器解耦本体感知、外部感知与动作,减少干扰
- 提出交互图表示与稀疏边注意力机制,精准建模人与人之间空间关系
- 首次实现端到端文本驱动的多人类机器人物理动作生成,适合智能体协同研究
人类形态智能体需模拟人类社交行为中的复杂协作。然而现有方法大多局限于单智能体场景,忽视了多智能体互动中物理上合理的相互作用。为此,我们提出InterAgent,首个基于文本驱动的物理仿真多人类机器人控制端到端框架。核心是引入具备多流模块的自回归扩散变换器,解耦本体感知、外部感知与动作,减轻跨模态干扰并促进协同。进一步提出新型交互图外部感知表示,显式捕捉关节间精细空间依赖关系,助力网络学习。同时设计基于稀疏边的注意力机制,动态剔除冗余连接,强化关键跨智能体空间关系,提升交互建模鲁棒性。大量实验表明,InterAgent持续优于多个强基线,达到当前最优性能,仅通过文本提示即可生成连贯、物理合理且语义准确的多智能体行为。代码与数据将公开,以推动后续研究。
原文摘要 · Abstract (English)
Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAgent, the first end-to-end framework for text-driven physics-based multi-agent humanoid control. At its core, we introduce an autoregressive diffusion transformer equipped with multi-stream blocks, which decouples proprioception, exteroception, and action to mitigate cross-modal interference while enabling synergistic coordination. We further propose a novel interaction graph exteroception representation that explicitly captures fine-grained joint-to-joint spatial dependencies to facilitate network learning. Additionally, within it we devise a sparse edge-based attention mechanism that dynamically prunes redundant connections and emphasizes critical inter-agent spatial relations, thereby enhancing the robustness of interaction modeling. Extensive experiments demonstrate that InterAgent consistently outperforms multiple strong baselines, achieving state-of-the-art performance. It enables producing coherent, physically plausible, and semantically faithful multi-agent behaviors from only text prompts. Our code and data will be released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。