实时音频驱动角色动作,实现精准对齐与稳定长时序生成
DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

- 用因果音乐编码器捕捉节奏与相位动态,结合扩散强迫机制建模时序
- 在非平稳音频下保持响应速度与长期一致性,长时滚动更稳定
- 适用于交互式系统,支持实时虚拟人播放与人体机器人部署
我们研究实时音频驱动角色控制这一符合实际部署需求的问题:严格因果、低延迟流式处理,需在交互帧率下生成连贯的全身动作,且音频条件可能突变(如节奏变化、中断或用户编辑)。现有音乐到动作系统多针对离线生成优化,缺乏全局上下文,在流式推理中因条件历史过时而性能下降。我们提出DiscoForcing,一种流式音频驱动的扩散框架,结合因果音乐编码器以捕捉节奏结构和相位动态,并采用在时间跨度上异质噪声水平下训练的扩散强迫序列模型。在此基础上,设计混合时序调度与历史引导的流式采样器,显式权衡响应速度与非平稳音频下的长期一致性。在端到端实时交互系统中实现在线虚拟人播放与人形机器人部署,DiscoForcing 在匹配因果性与延迟约束下,相比基线展现出更稳定的长时滚动与更精确的音画对齐,同时维持实时吞吐能力。
原文摘要 · Abstract (English)
We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate coherent full-body motion at interactive frame rates while the audio condition can change abruptly, including tempo shifts, drops, or user edits. Prior music-to-motion systems are largely optimized for offline generation with global context, and degrade in streaming rollouts where conditioning history becomes stale or unreliable. We introduce DiscoForcing, a streaming audio-driven diffusion framework that combines a causal music encoder that captures rhythmic structure and phase dynamics with a diffusion-forcing sequence model trained under heterogeneous noise levels across the temporal horizon. Building on this, we design a hybrid temporal schedule and a history-guided streaming sampler to explicitly trade off responsiveness against long-horizon consistency under non-stationary audio. Implemented in an end-to-end real-time interactive system with online avatar playback and humanoid deployment workflows, DiscoForcing delivers more stable long-horizon rollouts and sharper audio-motion alignment than prior baselines under matched causality and latency constraints while maintaining real-time throughput.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。