arXiv:2505.06985cs.CV2025-05

让视频生成更像指定人物,通过自动传播关键特征提升一致性

BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation

  • 用自回归方式逐帧注入参考主体的结构纹理特征
  • 在CLIP-I和DINO指标上分别提升7.8和13.1分
  • 适合需要高保真人物一致性的视频生成场景

零样本和基于微调的定制化文生图(CT2I)在故事创作中已取得显著进展,但定制化文生视频(CT2V)研究仍较有限。现有零样本CT2V方法泛化能力差,而将微调式文生图模型与时间运动模块结合的方法常导致结构和纹理信息丢失。为此,我们提出自回归结构与纹理传播模块(STPM),从参考主体提取关键结构和纹理特征,并逐帧自动注入视频各帧以增强一致性。此外,引入测试时奖励优化(TTRO)进一步优化细粒度细节。定量与定性实验验证了STPM与TTRO的有效性,在基线基础上使CLIP-I与DINO一致性指标分别提升7.8和13.1分。

原文摘要 · Abstract (English)

Both zero-shot and tuning-based customized text-to-image (CT2I) generation have made significant progress for storytelling content creation. In contrast, research on customized text-to-video (CT2V) generation remains relatively limited. Existing zero-shot CT2V methods suffer from poor generalization, while another line of work directly combining tuning-based T2I models with temporal motion modules often leads to the loss of structural and texture information. To bridge this gap, we propose an autoregressive structure and texture propagation module (STPM), which extracts key structural and texture features from the reference subject and injects them autoregressively into each video frame to enhance consistency. Additionally, we introduce a test-time reward optimization (TTRO) method to further refine fine-grained details. Quantitative and qualitative experiments validate the effectiveness of STPM and TTRO, demonstrating improvements of 7.8 and 13.1 in CLIP-I and DINO consistency metrics over the baseline, respectively.

文生视频身份一致自回归生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。