arXiv:2512.03590cs.CV2025-12

让说话人脸视频在两帧间自然过渡,更真实还原口型动作。

Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling

论文配图:Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling
图 1 · 摘自论文原文
  • 用上下文感知建模,从语音和周边画面中恢复细微表情变化
  • 在Hallo3数据集上FID提升23.3%,FVD提升36.5%,效果领先
  • 适合需要精准控制人脸视频中间帧的影视剪辑与虚拟人应用

现有说话人脸生成方法多用于开放式生成,而非连接两个已有视频段。本文研究说话人脸插值这一实用编辑任务,旨在固定起止帧条件下生成逼真的中间帧。与通用视频插值不同,该任务需在长时序间隙中恢复微妙的言语驱动面部动态,仅靠边界帧难以提供充分指导。为此,我们提出BBF(Beyond Boundary Frames)——一种统一的上下文感知框架。其包含三个互补组件:端点锚定以保持边界一致性,运动演化建模利用周围视觉上下文捕捉合理的时间过渡,语音动态精修通过语音音频注入细粒度言语驱动的面部动作。采用渐进优化策略,在去噪过程中平衡结构一致性和运动精细化。在说话人脸基准HDTF与Hallo3上的大量实验表明,BBF始终达到当前最优性能。尤其在Hallo3上,相比最强基线,FID降低23.3%,FVD降低36.5%。此外,BBF在通用视频插值基准上也展现出强泛化能力。

原文摘要 · Abstract (English)

Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we study talking-head inbetweening, a practical editing task that aims to generate realistic intermediate frames under fixed endpoint constraints. Unlike generic video inbetweening, this task requires recovering subtle speech-driven facial dynamics over long temporal gaps, where the boundary frames alone provide insufficient guidance for realistic motion recovery. To address this problem, we propose BBF (Beyond Boundary Frames), a unified context-aware framework for talking-head inbetweening. BBF consists of three complementary components: Endpoint Anchoring for preserving endpoint consistency, Motion Evolution Modeling for capturing plausible temporal transitions from surrounding visual context, and Speech Dynamics Refinement for injecting fine-grained speech-driven facial dynamics from speech audio. A progressive optimization strategy further balances structural consistency and motion refinement during denoising. Extensive experiments on the talking-head benchmarks HDTF and Hallo3 demonstrate that BBF consistently achieves state-of-the-art performance. In particular, BBF surpasses the strongest baseline on Hallo3 by 23.3% in FID and 36.5% in FVD. Moreover, BBF demonstrates strong generalization on generic video inbetweening benchmarks.

说话人脸视频插值语音驱动动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。