arXiv:2409.15259cs.CVcs.AI2024-09被引 4

用空间与语法引导注意力,提升多对象视频生成语义对齐

StarVid: Enhancing Semantic Alignment in Video Diffusion Models via Spatial and SynTactic Guided Attention Refocusing

  • 通过大模型规划运动轨迹,提供空间先验指导注意力聚焦
  • 引入语法感知对比约束,强化动词与名词的对应关系
  • 无需训练、可即插即用,适合复杂场景视频生成任务

基于扩散模型的文本到视频生成近期取得显著进展,但在包含多个对象和独立运动的复合场景中,难以准确反映文本提示的语义。为此,我们提出星形视频(StarVid),一种即插即用、无需训练的方法,用于增强多主体及其运动与文本提示之间的语义对齐。StarVid首先利用大语言模型(LLM)的空间推理能力,根据文本提示进行两阶段运动轨迹规划,生成的空间先验引导空间感知损失,使跨注意力(CA)映射聚焦于特定区域。此外,提出语法引导的对比约束,加强动词对应名词的CA映射关联,提升运动-主体绑定效果。定性与定量评估表明,该框架显著优于基线方法,生成视频质量更高且语义一致性更强。

原文摘要 · Abstract (English)

Recent advances in text-to-video (T2V) generation with diffusion models have garnered significant attention. However, they typically perform well in scenes with a single object and motion, struggling in compositional scenarios with multiple objects and distinct motions to accurately reflect the semantic content of text prompts. To address these challenges, we propose \textbf{StarVid}, a plug-and-play, training-free method that improves semantic alignment between multiple subjects, their motions, and text prompts in T2V models. StarVid first leverages the spatial reasoning capabilities of large language models (LLMs) for two-stage motion trajectory planning based on text prompts. Such trajectories serve as spatial priors, guiding a spatial-aware loss to refocus cross-attention (CA) maps into distinctive regions. Furthermore, we propose a syntax-guided contrastive constraint to strengthen the correlation between the CA maps of verbs and their corresponding nouns, enhancing motion-subject binding. Both qualitative and quantitative evaluations demonstrate that the proposed framework significantly outperforms baseline methods, delivering videos of higher quality with improved semantic consistency.

视频生成扩散模型注意力机制语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。