arXiv:2510.21696cs.CV2025-10被引 1

无需参考图或训练,实现角色与背景一致的视频生成。

BachVid: Training-Free Video Generation with Consistent Background and Character

  • 通过分析DiT注意力机制,提取前景掩码和匹配点
  • 缓存身份视频中间变量,注入新视频保持一致性
  • 训练免费,适合需多视频一致性的创作场景

扩散Transformer(DiTs)近期推动了文本到视频(T2V)生成的重大进展。然而,生成多个角色和背景一致的视频仍是重大挑战。现有方法通常依赖参考图像或大量训练,且大多仅解决角色一致性,背景一致性仍由图像到视频模型处理。我们提出BachVid,首个无需训练且不依赖参考图像的一致性视频生成方法。该方法基于对DiT注意力机制和中间特征的系统分析,发现其在去噪过程中可提取前景掩码并识别匹配点。我们的方法首先生成身份视频并缓存中间变量,随后将这些变量注入新生成视频的对应位置,确保多个视频中前景与背景的一致性。实验表明,BachVid在无需额外训练的情况下实现了稳健的一致性,为无需参考图像或额外训练的一致性视频生成提供了新颖高效的解决方案。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) have recently driven significant progress in text-to-video (T2V) generation. However, generating multiple videos with consistent characters and backgrounds remains a significant challenge. Existing methods typically rely on reference images or extensive training, and often only address character consistency, leaving background consistency to image-to-video models. We introduce BachVid, the first training-free method that achieves consistent video generation without needing any reference images. Our approach is based on a systematic analysis of DiT's attention mechanism and intermediate features, revealing its ability to extract foreground masks and identify matching points during the denoising process. Our method leverages this finding by first generating an identity video and caching the intermediate variables, and then inject these cached variables into corresponding positions in newly generated videos, ensuring both foreground and background consistency across multiple videos. Experimental results demonstrate that BachVid achieves robust consistency in generated videos without requiring additional training, offering a novel and efficient solution for consistent video generation without relying on reference images or additional training.

视频生成一致性扩散模型免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。