arXiv:2608.25734cs.CV2026-08

让语音生成手势更精准可控,支持实时流式调整动作细节。

InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

论文配图:InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control
图 1 · 摘自论文原文
  • 通过可微分解码器在采样时动态引导动作潜空间,实现关节级空间控制。
  • 在BEAT2数据集上提升多关节定位精度,且保持手势整体自然度。
  • 适合需要精细手势控制的虚拟人、动画师或交互系统开发者使用。

语音伴随手势生成已取得显著进展,能从语音生成逼真的全身动作,但现有模型缺乏对单个关节的细粒度空间控制能力。为此,我们提出一种模型无关、推理时可用的 extit{InteractGesture}方法,通过可微分的RVQ-VAE解码器引导扩散采样器的目标潜变量,并反向传播空间控制梯度以调整采样过程中的运动潜变量。流式生成中的主要挑战是块间依赖性:标准顺序推理会冻结先前块,使后续块的空间约束无法回传调整已有轨迹,导致边界不一致。为此,我们提出 extit{Progressive Chunk Guidance},采用带延迟窗口的可编辑块集合策略,使空间约束能在流式生成中跨块反向传播梯度。在BEAT2数据集上的实验表明, extit{InteractGesture}在保持整体手势质量的同时,显著提升了多关节空间控制能力。该方法还可支持稀疏关节定位、密集关节轨迹控制及定向指指点等多样化应用。

原文摘要 · Abstract (English)

Co-speech gesture generation has made significant progress toward realistic full-body motion from speaker audio, yet existing models lack fine-grained spatial controllability of individual joints. To address this, we introduce \emph{InteractGesture}, a model-agnostic, inference-time method for spatially controllable gesture generation. \emph{InteractGesture} guides target latent estimates of a diffusion sampler through a differentiable RVQ-VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. A primary challenge in streaming co-speech generation is chunk-wise dependency: standard sequential inference freezes prior chunks, preventing spatial constraints in future chunks from adjusting preceding trajectories and causing boundary inconsistencies. To overcome this limitation, we propose \emph{Progressive Chunk Guidance}, a chunk-window strategy that maintains an active set of editable chunk latents with staggered delays, enabling spatial constraints to propagate gradients backward across chunk boundaries during streaming generation. Experiments on the BEAT2 dataset show that \emph{InteractGesture} improves multi-joint spatial control while preserving overall gesture quality. Furthermore, our approach supports diverse applications, including sparse joint positioning, dense joint trajectory control, and directional pointing. Our project page is available at https://exitudio.github.io/interactgesture-page .

手势生成扩散模型实时控制空间引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。