arXiv:2512.07328cs.CVcs.AI2025-12被引 5

让视频角色保持一致,连发型衣着都不变。

ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation

  • 用参考图联合重建与生成,全量感知角色信息。
  • 新注意力模块防止帧间身份漂移,提升一致性。
  • 适合需要角色连贯性的影视动画与虚拟人场景。

文本到视频(T2V)生成进展迅速,但保持角色跨场景的一致性仍是难题。现有个性化方法多关注面部特征,忽视发型、服装、体型等关键上下文线索,影响视觉连贯性。本文提出 ContextAnyone,一种上下文感知的扩散框架,仅需单张参考图即可实现角色一致的文本到视频生成。该方法联合重建参考图并生成新视频帧,使模型充分感知和利用参考信息。通过创新的强调注意力模块,有选择地强化参考相关特征,防止帧间身份漂移。双引导损失结合扩散与参考重建目标,提升外观保真度;提出的 Gap-RoPE 位置嵌入分离参考与视频令牌,稳定时序建模。实验表明,ContextAnyone 在身份一致性和视觉质量上优于现有参考图像到视频方法,能生成多样运动与场景下连贯且保留上下文的角色视频。

原文摘要 · Abstract (English)

Text-to-video (T2V) generation has advanced rapidly, yet maintaining consistent character identities across scenes remains a major challenge. Existing personalization methods often focus on facial identity but fail to preserve broader contextual cues such as hairstyle, outfit, and body shape, which are critical for visual coherence. We propose \textbf{ContextAnyone}, a context-aware diffusion framework that achieves character-consistent video generation from text and a single reference image. Our method jointly reconstructs the reference image and generates new video frames, enabling the model to fully perceive and utilize reference information. Reference information is effectively integrated into a DiT-based diffusion backbone through a novel Emphasize-Attention module that selectively reinforces reference-aware features and prevents identity drift across frames. A dual-guidance loss combines diffusion and reference reconstruction objectives to enhance appearance fidelity, while the proposed Gap-RoPE positional embedding separates reference and video tokens to stabilize temporal modeling. Experiments demonstrate that ContextAnyone outperforms existing reference-to-video methods in identity consistency and visual quality, generating coherent and context-preserving character videos across diverse motions and scenes. Project page: \href{https://github.com/ziyang1106/ContextAnyone}{https://github.com/ziyang1106/ContextAnyone}.

文本生成视频角色一致性扩散模型上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。