arXiv:2502.06527cs.CV2025-02被引 13

用参考图生成个性视频,解决动态一致性难题

CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers

  • 引入3D参考注意力,让参考图与所有视频帧实时联动
  • 提出时间感知偏置策略,避免参考图过度影响生成结果
  • 适合需要个性化视频生成的创作者与研究人员

个性化视频生成因时序不一致和质量下降仍具挑战。本文提出CustomVideoX框架,基于视频扩散变换器,仅通过训练LoRA参数提取参考图像特征,实现高效自适应。为实现参考图与视频内容的无缝交互,提出3D参考注意力机制,使参考特征在空间与时间维度上同时作用于所有视频帧。为缓解推理时参考图特征与文本引导的过度影响,设计时间感知参考注意力偏置(TAB)策略,动态调节不同时间步的参考偏置。此外,引入实体区域感知增强(ERAE)模块,通过调整注意力偏置,将关键实体标记的高激活区域与参考特征注入对齐。为全面评估个性化视频生成,构建新基准VideoBench,包含超过50个物体和100条提示。实验表明,CustomVideoX在视频一致性和质量上显著优于现有方法。

原文摘要 · Abstract (English)

Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal inconsistencies and quality degradation. In this paper, we introduce CustomVideoX, an innovative framework leveraging the video diffusion transformer for personalized video generation from a reference image. CustomVideoX capitalizes on pre-trained video networks by exclusively training the LoRA parameters to extract reference features, ensuring both efficiency and adaptability. To facilitate seamless interaction between the reference image and video content, we propose 3D Reference Attention, which enables direct and simultaneous engagement of reference image features with all video frames across spatial and temporal dimensions. To mitigate the excessive influence of reference image features and textual guidance on generated video content during inference, we implement the Time-Aware Reference Attention Bias (TAB) strategy, dynamically modulating reference bias over different time steps. Additionally, we introduce the Entity Region-Aware Enhancement (ERAE) module, aligning highly activated regions of key entity tokens with reference feature injection by adjusting attention bias. To thoroughly evaluate personalized video generation, we establish a new benchmark, VideoBench, comprising over 50 objects and 100 prompts for extensive assessment. Experimental results show that CustomVideoX significantly outperforms existing methods in terms of video consistency and quality.

视频生成扩散模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。