arXiv:2607.08741cs.GRcs.CV2026-07被引 2

实时交互式生成高精度3D人体动作,支持文本与姿态约束控制

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

论文配图:ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
图 1 · 摘自论文原文
  • 采用根节点显式特征与隐变量身体嵌入的混合表示
  • 两阶段自回归去噪器支持长时序约束与在线文本提示
  • 可在真实交互场景中实现路径跟随与鼠标键盘控制

在动画、仿真和人形机器人等交互应用中,实现实时生成逼真3D人体动作至关重要。现有离线方法虽能精确控制,但推理速度不足;而在线方法虽支持实时生成,却常牺牲可控性或难以处理复杂语义和长时序目标。本文提出ARDY,一种流式生成框架,通过在线文本提示与灵活的运动学约束实现高质量动作生成。该框架采用显式根部特征与潜在身体嵌入的混合表示,在轨迹控制精度与生成效率间取得平衡。设计了具备可变历史上下文的两阶段自回归变压器去噪器,支持对长时序运动学约束的条件生成。基于大规模动作捕捉数据集训练,并直接以真实动作中的文本标签和运动学约束作为条件,模型天然具备可控生成能力。在HumanML3D基准与高保真Bones Rigplay数据集上的评估显示,ARDY在动作质量与约束遵循性上表现优异。最后,通过交互演示验证其灵活性:支持动态文本控制、多种关键帧约束、路径跟随及鼠标键盘交互式行走控制。补充视频、代码与模型已发布于https://research.nvidia.com/labs/sil/projects/ardy/。

原文摘要 · Abstract (English)

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We propose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method's practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at https://research.nvidia.com/labs/sil/projects/ardy/.

动作生成扩散模型交互控制自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。