arXiv:2603.13500cs.CV2026-03被引 4

用帧级动作规划实现实时流式生成与高质量离线生成统一

ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning

  • 每帧预测文本隐变量作为语义锚点,结合运动线索逐步去噪
  • 实时流式速度提升5.25倍,FID降低18%且质量优于现有方法
  • 支持零样本动作编辑和补帧,无需额外模型

我们提出ActionPlan,一种统一的运动扩散框架,将实时流式生成与高质量离线生成统一于单一模型中。核心思想是引入帧级动作计划:模型预测每帧的文本隐变量,作为去噪过程中的密集语义锚点,并利用语义与运动线索联合去噪完整动作序列。为支持此结构化流程,设计了针对隐变量的扩散步骤,使每个运动隐变量可独立去噪,并在推理时灵活采样顺序。结果表明,ActionPlan可在历史条件、未来感知模式下实现实时流式生成,同时支持高质量离线生成。相同机制还实现零样本动作编辑与补帧,无需额外模型。实验显示,其实时流式速度比最优先前方法快5.25倍,且在FID指标上提升18%,运动质量更优。

原文摘要 · Abstract (English)

We present ActionPlan, a unified motion diffusion framework that bridges real-time streaming with high-quality offline generation within a single model. The core idea is to introduce a per-frame action plan: the model predicts frame-level text latents that act as dense semantic anchors throughout denoising, and uses them to denoise the full motion sequence with combined semantic and motion cues. To support this structured workflow, we design latent-specific diffusion steps, allowing each motion latent to be denoised independently and sampled in flexible orders at inference. As a result, ActionPlan can run in a history-conditioned, future-aware mode for real-time streaming, while also supporting high-quality offline generation. The same mechanism further enables zero-shot motion editing and in-betweening without additional models. Experiments demonstrate that our real-time streaming is 5.25x faster while also achieving 18% motion quality improvement over the best previous method in terms of FID.

动作生成扩散模型实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。