用未来关键帧引导,让语音驱动的人体动画更像自己。
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
- 用未来时间步的关键帧作指引,而非当前窗口内固定锚点。
- 在三个模型上提升唇音同步与身份一致性,视觉质量更好。
- 可自动生成参考图像为锚点,无需额外生成关键帧。
语音驱动的人体动画模型在时序自回归生成中常出现身份漂移问题,导致角色随时间逐渐失真。现有方法通过生成中间关键帧作为时间锚点来缓解,但需额外生成阶段且限制自然动作。本文提出前瞻锚定(Lookahead Anchoring),利用当前生成窗口之外的未来时间步关键帧作为动态指引,将关键帧从固定边界转化为方向性参照物。模型在响应即时音频的同时持续向未来锚点靠近,实现身份持久保持。该方法还支持自键帧机制:以参考图作为前瞻目标,完全省去关键帧生成。我们发现,前瞻距离自然调节表达力与一致性间的平衡——距离越大,动作越自由;距离越小,身份越稳定。在三个近期人体动画模型上应用后,该方法显著提升了唇音同步、身份保真度与视觉质量,证明其在多种架构中均具备优越的时序建模能力。视频结果见:https://lookahead-anchoring.github.io。
原文摘要 · Abstract (English)
Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermediate temporal anchors that prevent degradation, but this requires an additional keyframe generation stage and can restrict natural motion dynamics. To address this, we propose Lookahead Anchoring, which leverages keyframes from future timesteps ahead of the current generation window, rather than within it. This transforms keyframes from fixed boundaries into directional beacons: the model continuously pursues these future anchors while responding to immediate audio cues, maintaining consistent identity through persistent guidance. This also enables self-keyframing, where the reference image serves as the lookahead target, eliminating the need for keyframe generation entirely. We find that the temporal lookahead distance naturally controls the balance between expressivity and consistency: larger distances allow for greater motion freedom, while smaller ones strengthen identity adherence. When applied to three recent human animation models, Lookahead Anchoring achieves superior lip synchronization, identity preservation, and visual quality, demonstrating improved temporal conditioning across several different architectures. Video results are available at the following link: https://lookahead-anchoring.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。