arXiv:2607.17985cs.CVcs.MM2026-07

让视频主角按指令连贯动作且不跑偏,靠关键帧锚定身份。

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

论文配图:Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation
图 1 · 摘自论文原文
  • 用关键帧分段生成,结合前后帧保持主体外观一致。
  • 在官方评测中排名第三,支持多动作序列连续生成。
  • 适合需要角色一致性视频生成的创作者或研究者。

身份保持的文本到视频生成旨在根据文本描述合成视频,同时确保用户指定主体在整个过程中保持可识别性。IPVG26挑战将该框架从单一整体提示扩展到时间结构化描述:模型接收一系列带时间戳的动作字幕,需按顺序生成主体执行这些动作的视频。这种时间结构带来新挑战——主体需持续完成一系列不同动作,同时保持身份一致。然而,端到端视频生成器在运动累积和动作变化下易出现外观漂移。为此,我们提出无需训练的三阶段流水线:首先通过动作感知提示优化,将输入重写为指定每个动作终点状态的图像生成提示;其次通过身份保持生成阶段,联合参考身份与前一帧生成关键帧序列,实现外观与姿态解耦;最后通过多参考引导与身份驱动噪声搜索增强中间片段合成,强化采样过程中的身份保真度。本方法在官方Track 2排行榜中位列第三,展现出良好性能与强泛化能力。

原文摘要 · Abstract (English)

Identity-preserving text-to-video generation aims to synthesize a video that accurately follows a textual description while maintaining the recognizability of a user-specified subject throughout. The IPVG26 challenge extends this framework from a single holistic prompt to a temporally structured specification. The model additionally receives a sequence of timestamped action captions and must render the subject performing these actions in the specified order. This temporal structure presents a challenge not encountered in previous identity-preserving generation tasks, as the subject must continuously perform a scripted sequence of distinct actions while maintaining a consistent identity. However, end-to-end video generators are prone to appearance drift as motion accumulates and the depicted actions change. We address this challenge with a training-free, three-stage pipeline framework. An action-aware prompt polishment stage first rewrites the inputs into image-generation prompts that specify the terminal state of each action. An identity-preserving generation stage then produces the keyframe sequence by conditioning each frame jointly on the reference identity and its predecessor, thereby decoupling time-invariant appearance from time-varying pose. Finally, an identity-aware inference enhancement stage synthesizes the intermediate segments using multi-reference guidance and identity-driven noise searching, both of which reinforce identity fidelity during sampling. Our method ranked third on the official Track 2 leaderboard, demonstrating competitive performance and strong generality.

视频生成身份保持关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。