arXiv:2512.19546cs.CV2025-12被引 7

让虚拟人说话动作精准跟随文本,时间对齐更自然。

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

  • 分阶段注意力机制,让模型理解动作时序和语义。
  • 在多个数据集上优于现有方法,动作控制精度显著提升。
  • 无需额外姿态信号,适合影视动画与虚拟主播场景。

尽管说话头像生成已取得显著进展,现有方法仍面临三大挑战:动作对文本的跟随能力不足、动作与音频内容缺乏时间对齐、依赖额外控制信号(如姿态骨骼)。我们提出 ActAvatar 框架,通过文本引导实现动作控制的相位级精确性,同时捕捉动作语义与时间上下文。核心创新包括:(1) 相位感知交叉注意力(PACA),将提示分解为全局基础块与时间锚定的相位块,使模型聚焦于相关时段的关键词;(2) 渐进式音视频对齐,早期层优先利用文本构建动作结构,深层则强化音频以精细调整口型,避免模态干扰;(3) 两阶段训练策略:先在多样数据上建立强音视频对应关系,再通过结构化标注微调注入动作控制能力,兼顾音视频对齐与文本跟随性能。大量实验表明,ActAvatar 在动作控制与视觉质量方面均显著超越当前最优方法。

原文摘要 · Abstract (English)

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We present ActAvatar, a framework that achieves phase-level precision in action control through textual guidance by capturing both action semantics and temporal context. Our approach introduces three core innovations: (1) Phase-Aware Cross-Attention (PACA), which decomposes prompts into a global base block and temporally-anchored phase blocks, enabling the model to concentrate on phase-relevant tokens for precise temporal-semantic alignment; (2) Progressive Audio-Visual Alignment, which aligns modality influence with the hierarchical feature learning process-early layers prioritize text for establishing action structure while deeper layers emphasize audio for refining lip movements, preventing modality interference; (3) A two-stage training strategy that first establishes robust audio-visual correspondence on diverse data, then injects action control through fine-tuning on structured annotations, maintaining both audio-visual alignment and the model's text-following capabilities. Extensive experiments demonstrate that ActAvatar significantly outperforms state-of-the-art methods in both action control and visual quality.

虚拟人动作控制音视频对齐文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。