通过关键帧与文本引导,提升生成中间帧的语义一致性和运动稳定性。
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
- 用关键帧和文本构建注意力偏置,指导中间帧生成路径。
- 在短/长序列上均实现最佳帧一致性与运动稳定性,无需额外训练。
- 适用于需要精准控制动作流程的视频生成任务,如动画创作。
生成式插帧(GI)旨在生成首尾关键帧之间的真实感中间帧,而非简单插值。当序列间隔变大、运动幅度增加时,现有模型常出现帧间不一致、节奏不稳定和语义错位问题。由于固定端点且存在多种合理生成路径,该任务需依赖关键帧和文本提供额外引导以明确目标轨迹。为此,本文提出关键帧锚定注意力偏置(Keyframe-anchored Attention Bias),从关键帧与文本中为每帧中间帧注入语义与时间引导。同时引入重缩放时间位置编码(Rescaled Temporal RoPE),增强自注意力对关键帧的忠实建模能力。我们还构建了首个专用于文本条件插帧评估的基准测试TGI-Bench,支持针对性分析。无需额外训练,本方法在各类挑战下均达到当前最优的帧一致性、语义保真度与节奏稳定性,适用于短序列与长序列场景。
原文摘要 · Abstract (English)
Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent frames with unstable pacing and semantic misalignment. Since GI involves fixed endpoints and numerous plausible paths, this task requires additional guidance gained from the keyframes and text to specify the intended path. Thus, we give semantic and temporal guidance from the keyframes and text onto each intermediate frame through Keyframe-anchored Attention Bias. We also better enforce frame consistency with Rescaled Temporal RoPE, which allows self-attention to attend to keyframes more faithfully. TGI-Bench, the first benchmark specifically designed for text-conditioned GI evaluation, enables challenge-targeted evaluation to analyze GI models. Without additional training, our method achieves state-of-the-art frame consistency, semantic fidelity, and pace stability for both short and long sequences across diverse challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。