arXiv:2608.16154cs.CVcs.GR2026-08中稿 · ACM MM 2026

无需训练即可保持视频主体身份一致,支持复杂动作生成。

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

论文配图:KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation
图 1 · 摘自论文原文
  • 分离动态生成与身份注入,用关键帧修正身份
  • 在长序列动作中保持身份一致性,优于现有方法
  • 支持多主体参考,适合需要高保真身份的视频生成

身份保持视频生成(IPVG)需同时忠实于参考主体和文本提示。现有方法常因调优成本高或输入增强有限,难以在复杂长序列动作中维持严格的身份一致性。为此,我们提出KeyID,一种无需训练的IPVG框架,将视频动态合成与身份注入解耦。该框架包含两部分:(1) 参考感知视频生成,生成与多个参考对齐但不依赖身份的视频草稿;(2) 身份保留关键帧编辑,通过稀疏关键帧修正并插值运动实现身份注入。通过从密集帧监督转为稀疏关键帧优化,有效缓解了提示遵循与身份保真之间的容量冲突。其模块化设计可无缝扩展至多主体参考与复杂序列动作生成,无需额外训练。KeyID在官方挑战赛基准上经自动与人工评估验证,最终获得ACM Multimedia 2026 IPVG Grand Challenge Track 2(序列动作)亚军。源代码见https://github.com/WISLab-GDUT/KeyID。

原文摘要 · Abstract (English)

Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enhancements, struggling to maintain rigid identity consistency during complex, long-sequence actions. To address these limitations, we propose KeyID, a training-free IPVG framework that decouples the synthesis of video dynamics from the injection of identity. Specifically, KeyID comprises two components: (1) Reference-Aware Video Generation, which produces an identity-agnostic video draft aligned with multiple references, and (2) Identity-Preserved Keyframe Editing, which integrates the target identity via sparse keyframe correction and subsequent motion interpolation. By shifting from dense frame-level supervision to sparse keyframe-level refinement, KeyID effectively resolves the capacity conflict between prompt adherence and identity fidelity. Crucially, our modular design allows seamless extension to multi-subject references and complex sequential action generation without additional training. KeyID outperforms prior works and is validated by automatic and human evaluations on the official challenge benchmark, ultimately securing the runner-up position in the Track 2 (Sequential Action) of the ACM Multimedia 2026 IPVG Grand Challenge. Source code is available at https://github.com/WISLab-GDUT/KeyID.

视频生成身份保持关键帧编辑无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。