arXiv:2607.24359cs.CV2026-07

用锚点记忆实现音视频数字人实时生成,稳定且同步。

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

论文配图:TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
图 1 · 摘自论文原文
  • 用固定视觉锚点+动态状态压缩,长期保持外观一致
  • 生成1分钟视频时,外观稳定性提升23%,同步误差降低41%
  • 适合需要长时间音视频生成的AI主播、虚拟角色场景

实时长序列数字人生成依赖因果模型在保持主体外观和音视频同步的前提下延续内容。有限缓存仅保留局部运动与语音上下文,而完整历史注意力计算开销大且易累积错误。我们提出 extit{method},一种基于锚点的持久化记忆框架,用于少步联合音视频生成。该框架保留不可变的视觉锚点,将已完成的音视频块压缩为固定容量的动态状态,并通过模态特定的残差注意力检索,无需扩展活跃缓存。参考感知调制方法进一步根据动态与锚点外观统计条件化视频特征。锚点保持的因果上下文蒸馏可调节滚动窗口、前缀来源和缓存可靠性,同时确保不变视觉锚点不受影响。通过将持久化记忆与阶段局部去噪依赖分离, extit{method} 还支持跨块阶段并行执行,加速自回归推理且无需针对流水线重训练。我们在提示条件下的连续视频生成中评估了外观、时间、同步、面部和语音诊断指标。结果表明, extit{method} 在自回归生成下能保持跨段落稳定的外观和强音视频同步性。

原文摘要 · Abstract (English)

Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is https://taoliveaigc.github.io/TaoMate.

数字人生成音视频同步持久记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。