arXiv:2503.14428cs.CVcs.AI2025-03

让视频生成更准确地表现多个主体及其关系,解决内容缺失和位置错位问题。

Comp-Attn: Present-and-Align Attention for Compositional Video Generation

  • 通过条件插值与注意力调制,分离并分别处理主体存在与空间关系对齐问题。
  • 在多个基准上提升15.7%的组合视频生成分数,推理时间仅增加5%。
  • 无需训练即可适配多种模型,适合需要精确控制多主体交互的生成任务。

在文本到视频(T2V)生成领域,可靠合成涉及多个主体及复杂关系的组合内容仍处于探索阶段。主要挑战包括:1)主体存在性,部分主体无法在视频中出现;2)主体间关系,主体间的互动与空间关系对齐不准。现有方法如推理时潜在优化或布局控制,难以同时解决两个问题。为此,我们提出Comp-Attn,一种遵循“呈现-对齐”范式的组合感知交叉注意力变体:通过条件层面确保主体存在,注意力分布层面实现关系对齐。具体而言,1)引入主体感知条件插值(SCI),强化主体特定条件,确保每个主体的出现;2)提出布局强制注意力调制(LAM),动态引导注意力分布以匹配多主体的相对布局。Comp-Attn可无训练地集成至多种T2V基线,在Wan2.1-T2V-14B和Wan2.2-T2V-A14B上分别提升T2V-CompBench得分15.7%和11.7%,推理时间仅增加5%。同时在VBench和T2I-CompBench上表现优异,证明其在通用视频生成与组合文本到图像任务中的可扩展性。

原文摘要 · Abstract (English)

In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relations, where the interaction and spatial relationship between subjects are misaligned. Existing methods adopt techniques, such as inference-time latent optimization or layout control, which fail to address both issues simultaneously. To tackle these problems, we propose Comp-Attn, a composition-aware cross-attention variant that follows a Present-and-Align paradigm: it decouples the two challenges by enforcing subject presence at the condition level and achieving relational alignment at the attention-distribution level. Specifically, 1) We introduce Subject-aware Condition Interpolation (SCI) to reinforce subject-specific conditions and ensure each subject's presence; 2) We propose Layout-forcing Attention Modulation (LAM), which dynamically enforces the attention distribution to align with the relational layout of multiple subjects. Comp-Attn can be seamlessly integrated into various T2V baselines in a training-free manner, boosting T2V-CompBench scores by 15.7\% and 11.7\% on Wan2.1-T2V-14B and Wan2.2-T2V-A14B with only a 5\% increase in inference time. Meanwhile, it also achieves strong performance on VBench and T2I-CompBench, demonstrating its scalability in general video generation and compositional text-to-image (T2I) tasks.

视频生成组合生成注意力机制多主体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。