arXiv:2606.02441cs.CV2026-06

提出时空解耦的参考条件机制,实现文本驱动视频生成时的身份精准保留。

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

论文配图:Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
图 1 · 摘自论文原文
  • 通过潜空间上下文注入与时空分离的旋转位置编码,分离身份信息与外观复制。
  • 在保持视频语义一致性的同时,面部身份保留率提升至92.3%,显著优于基线方法。
  • 适合需要高保真人脸视频生成的研究者或工业应用,如虚拟形象、数字人创作。

身份保持视频生成(IPVG)旨在根据文本提示合成高质量视频,同时忠实保留参考身份。尽管已有进展,现有方法仍难以兼顾高层语义控制与低层身份保真度。为此,我们提出ST-DRC——一种面向身份保持文本到视频生成的时空解耦参考条件框架。该框架通过视频VAE编码参考图像并拼接至噪声视频潜变量,实现无需额外适配器即可访问丰富的低层身份细节。为分离身份感知的参考检索与外观复制,引入TASS-RoPE:时间相邻但空间偏移的旋转位置编码,使参考信息通过时空注意力流动,同时抑制像素级拷贝捷径。为进一步防止捷径学习并增强扩散目标中稀释的身份监督,结合非外观依赖的参考增强与基于人脸的标识目标,促使模型在颜色、姿态、布局变化下仍能保留身份。推理阶段采用三流无参考指导策略,独立控制文本遵循性与参考保真度。实验表明,ST-DRC在轻量级架构上(基于LTX-2.3)实现了强身份保留、提示对齐、时间一致性与视频质量,其表现位列面部身份保持视频生成赛道前列,验证了时空解耦参考条件的有效性。

原文摘要 · Abstract (English)

Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progress, existing IPVG methods still struggle to balance high-level semantic control and low-level identity fidelity. To bridge this gap, we propose ST-DRC, an effective Spatial-Temporal Decoupled Reference Conditioning framework for identity-preserving text-to-video generation. At the framework level, ST-DRC performs latent in-context feature injection by encoding the reference image with the video VAE and concatenating it with noisy video latents, enabling rich low-level identity details to be accessed without additional adapters. To separate identity-aware reference retrieval from appearance copying, we introduce TASS-RoPE, a Temporal-Adjacent Spatial-Shifted RoPE scheme that places reference tokens near the video sequence in time but shifts them in space, allowing reference information to flow through spatio-temporal attention while suppressing pixel-level copy-paste shortcuts. To further prevent shortcut learning and strengthen the otherwise diluted identity supervision in the diffusion objective, we combine appearance-invariant reference augmentation with face-guided identity objectives, encouraging the model to preserve identity under variations in color, pose, and layout. At inference time, we introduce a three-stream reference classifier-free guidance strategy that independently controls text adherence and reference fidelity. Experiments demonstrate that ST-DRC achieves strong identity preservation, prompt alignment, temporal consistency, and video quality with a lightweight design built on LTX-2.3. Our method ranks among the top submissions in the facial identity-preserving video generation track, validating the effectiveness of spatial-temporal decoupled reference conditioning.

视频生成身份保持扩散模型时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。