arXiv:2605.20209cs.GRcs.LG2026-05被引 1

用强化学习快速精准控制角色动作,无需迭代优化。

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

论文配图:NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
图 1 · 摘自论文原文
  • 通过强化学习直接预测任务优化的噪声,实现高效控制。
  • 在多个任务中成功率更高,推理速度比传统方法快。
  • 适合需要快速响应的物理动画场景,如游戏与虚拟仿真。

在物理驱动的动画中实现精确、多功能的全身角色控制仍具挑战性。近期基于扩散模型的策略虽能生成丰富多样的动作,但通常依赖梯度引导来满足任务目标,过程缓慢且易降低鲁棒性。本文提出NaP-Control(Navigating Diffusion Prior for Versatile and Fast Character Control),简称NaP。该方法利用强化学习操控一个任务无关的扩散策略先验的潜在噪声,将其引导至特定任务行为,实现快速、鲁棒且高保真的运动控制。与仅依赖离线训练的方法不同,NaP在训练期间与环境交互以修正动作并优化任务奖励,提升成功率,并可适应复杂场景。通过直接预测任务优化的扩散噪声,NaP省去了去噪过程中的迭代引导,实现高效推理。实验表明,NaP在多样任务中均取得更高成功率和更快推理速度,同时保持自然的动作质量。

原文摘要 · Abstract (English)

Achieving precise, versatile whole-body character control in physics-based animation remains challenging. Recent diffusion-based policies generate rich and expressive motions but typically rely on gradient-based test-time guidance to satisfy task objectives, which is slow and can reduce robustness. We introduce NaP-Control (Navigating Diffusion Prior for Versatile and Fast Character Control), abbreviated as NaP. Our method uses reinforcement learning to manipulate the latent noise of a task-agnostic diffusion policy prior, steering it toward task-specific behaviors for fast, robust control with high motion fidelity. In contrast to methods that rely solely on offline training, NaP interacts with the environment during training to correct motions and optimize task rewards, improving success rates and enabling adaptation to challenging scenarios. By directly predicting task-optimized diffusion noise, NaP eliminates iterative guidance during denoising and enables efficient inference. Experiments show that NaP attains higher success rates and faster inference while preserving natural motion across diverse tasks.

角色控制扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。