arXiv:2511.10076cs.CV2025-11AAAI被引 14

用全局旋转扩散解决说话动作生成中的误差累积问题

Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints

  • 首次在全局旋转空间进行扩散生成,打破关节依赖链
  • 多层级约束使动作更平滑准确,性能提升46.0%
  • 适合需要稳定长时说话动作的虚拟人与动画应用

可靠的长时序共话语音动作生成需要精确的动作表示和全关节结构一致性。现有生成方法通常基于骨架结构定义的局部关节旋转,导致生成过程中误差累积,末端执行器动作不稳定且不自然。本文提出GlobalDiff,首个直接在全局关节旋转空间运行的扩散框架,从根本上解耦各关节预测与上游依赖,缓解层级误差积累。为弥补全局旋转空间缺乏结构先验的问题,引入多层级约束:关节结构约束通过虚拟锚点捕捉精细朝向;骨架结构约束保持骨骼间角度一致性以维持结构完整;时间结构约束采用多尺度变分编码器对齐生成动作与真实时间模式。三重约束联合正则化全局扩散过程,增强结构感知。在标准共话语音基准上大量评估显示,GlobalDiff生成动作更流畅精准,相比当前最强方法在多种说话人身份下性能提升46.0%。

原文摘要 · Abstract (English)

Reliable long-horizon co-speech gesture generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically operate on local joint rotations, which are defined hierarchically based on the skeleton structure. This leads to cumulative errors during generation, manifesting as unstable and implausible motions at end-effectors. In this work, we propose GlobalDiff, a diffusion-based framework that operates directly in the space of global joint rotations for the first time, fundamentally decoupling each joint's prediction from upstream dependencies and alleviating hierarchical error accumulation. To compensate for the absence of structural priors in global rotation space, we introduce a multi-level constraint scheme. Specifically, a joint structure constraint introduces virtual anchor points around each joint to better capture fine-grained orientation. A skeleton structure constraint enforces angular consistency across bones to maintain structural integrity. A temporal structure constraint utilizes a multi-scale variational encoder to align the generated motion with ground-truth temporal patterns. These constraints jointly regularize the global diffusion process and reinforce structural awareness. Extensive evaluations on standard co-speech benchmarks show that GlobalDiff generates smooth and accurate motions, improving the performance by 46.0% compared to the current SOTA under multiple speaker identities.

动作生成扩散模型语音驱动全局优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。