用检索引导噪声优化,实现复杂动作约束下的精准生成
Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization

- 从大数据集检索参考动作,引导扩散模型生成
- 在严苛时空约束下(如步数、障碍物)成功生成合理动作
- 结合大模型自动推理该检索什么,适合虚拟角色控制
生成满足自定义零样本目标函数的人体运动,可应用于可控角色动画和虚拟代理行为合成,是关键能力。现有方法虽能处理多种未见约束,但在面对严重时空限制的任务(如密集空间障碍或指定步数)时表现不佳。为应对这些高约束任务,我们提出一种基于无训练扩散噪声优化框架的检索引导方法。核心思想是从大规模动作数据集中搜索可能满足困难约束的参考。引入关系任务解析,将目标约束分组并识别需通过检索参考解决的难点。通过奖励引导掩码,将随机噪声与检索噪声结合,获得更优的扩散噪声初始值。在此基础上优化扩散噪声,成功解决高约束生成任务。借助大语言模型进行关系任务解析,整个框架可自动推理应检索的内容,提升了无训练优化方案下移动代理的智能性。
原文摘要 · Abstract (English)
Generating human motion that satisfies customized zero-shot goal functions, enabling applications such as controllable character animation and behavior synthesis for virtual agents, is a critical capability. While current approaches handle many unseen constraints, they fail on tasks with very challenging spatiotemporal restrictions, such as severe spatial obstacles or specified numbers of walking steps. To equip motion generators for these highly constrained tasks, we present a retrieval-guided method built on the training-free diffusion noise optimization framework. The key idea is to search within large motion datasets for guidance that can potentially satisfy difficult constraints. We introduce relational task parsing to group target constraints and identify the difficult ones to be handled by retrieved reference. A better initialization for diffusion noise is then obtained via a reward-guided mask that combines random noise with retrieved noise. By optimizing diffusion noise from this improved initialization, we successfully solve highly constrained generation tasks. By leveraging LLM for relational task parsing, the whole framework is further enabled to automatically reason for what to retrieve, improving the intelligence of moving agents under a training-free optimization scheme.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。