用锚帧保持角色视频长时间一致性,效果优于现有方法。
Gloria: Consistent Character Video Generation via Content Anchors

- 用少量锚帧捕捉角色视觉特征,稳定生成一致性外观。
- 支持超过10分钟长视频生成,多视角外观与身份高度一致。
- 适合影视动画、虚拟人开发等需要长期角色连贯性的场景。
数字角色在现代媒体中至关重要,但生成长时间、多视角一致且具表现力的角色视频仍具挑战。现有方法或缺乏足够上下文维持身份,或依赖非角色中心信息作为记忆,导致一致性不足。本文提出通过一组紧凑的锚帧表征角色视觉属性,提供稳定参考。针对基于参考的生成易出现复制粘贴和多锚冲突的问题,引入超集内容锚定机制,利用训练内外片段提示防止重复;并采用旋转位置编码作为弱条件,编码位置偏移以区分多个锚帧。此外,构建可扩展的流水线从海量视频中提取锚帧。实验表明,该方法可生成超过10分钟的高质量角色视频,在跨视角外观与身份表达上均优于现有方法。
原文摘要 · Abstract (English)
Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information as the memory, leading to suboptimal consistency. Recognizing that character video generation inherently resembles an outside-looking-in scenario. In this work, we propose representing the character visual attributes through a compact set of anchor frames. This design provides stable references for consistency, while reference-based video generation inherently faces challenges of copy-pasting and multi-reference conflicts. To address these, we introduce two mechanisms: Superset Content Anchoring, providing intra- and extra-training clip cues to prevent duplication, and RoPE as Weak Condition, encoding positional offsets to distinguish multiple anchors. Furthermore, we construct a scalable pipeline to extract these anchors from massive videos. Experiments show our method generates high-quality character videos exceeding 10 minutes, and achieves expressive identity and appearance consistency across views, surpassing existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。