arXiv:2607.04311cs.CV2026-07

让多人视频生成更真实一致,靠AI导演级描述和视觉语言模型对齐。

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

论文配图:Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
图 1 · 摘自论文原文
  • 用AI导演级文本描述+视觉语言模型提取多模态特征。
  • 在多个参考主体下保持身份一致,生成质量优于现有方法。
  • 适合需要复杂角色互动的视频生成研究者或创作者。

主体驱动与多元素视频生成是可控视频合成的核心挑战,现有方法仍难以保持身份一致性并建模多个主体间的复杂关系。本文提出Aura,一个统一的高保真、身份一致的视频生成框架。为更好捕捉场景动态与主体交互,引入AI导演级字幕,提供密集且结构化的视频内容描述。进一步利用带可学习查询的视觉语言模型(VLM),从文本与视觉参考中提取涵盖全局语义与细粒度视觉线索的多模态特征。为弥合VLM与扩散变换器(DiT)之间的表征差距,设计两阶段对齐策略,逐步将VLM特征映射至DiT特征空间。视觉条件输入采用标记拼接方式直接注入生成过程。为区分异构主体类型并减少复制粘贴伪影,提出主体感知的RoPE-Shift机制;为更好区分不同类别的参考图像,引入主体感知可学习标记。此外,引入记忆标记以平衡不同参考主体数量样本的训练信号。推理时,采用渐进式自适应提示引导(Progressive-APG)缓解过饱和问题,提升与用户提示的语义对齐。最后,通过专用数据构建流程建立高质量视频-主体图像数据集。大量实验表明,该方法在单主体生成及更具挑战性的多元素场景中均达到当前最优性能。

原文摘要 · Abstract (English)

Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model complex relationships among multiple subjects. In this paper, we propose Aura, a unified framework for high-fidelity and identity-consistent video generation. To better capture scene dynamics and subject interactions, we introduce AI director-level captions that provide dense and structured descriptions of video content. We further leverage a vision-language model (VLM) with learnable queries to extract multimodal semantic features from textual and visual references, covering both global semantics and fine-grained visual cues. To bridge the representational gap between the VLM and the Diffusion Transformer (DiT), we design a two-stage alignment strategy that progressively maps VLM features into the DiT feature space. For visual conditioning, we adopt token concatenation to inject reference information directly into the generation process. To distinguish heterogeneous subject types and reduce common copy-paste artifacts, we develop a subject-aware RoPE-Shift mechanism. To further differentiate reference images of different categories, we introduce subject-aware learnable tokens. In addition, we introduce Memory Tokens to balance the training signal across examples with different numbers of reference subjects. During inference, Progressive-APG (Adaptive Prompt Guidance) further alleviates oversaturation and improves semantic alignment with user prompts. Finally, we build a high-quality video-subject image dataset through a dedicated data construction pipeline. Extensive experiments show that our method achieves state-of-the-art performance on both single-subject generation and more challenging multi-element scenarios.

视频生成多主体一致性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。