arXiv:2505.22980cs.CV2025-05被引 3

无需训练即可精准生成多物体动态视频,提升复杂交互效果。

MOVi: Training-free Text-conditioned Multi-Object Video Generation

  • 用大语言模型当导演规划物体运动轨迹,通过重初始化噪声控制动作。
  • 在现有模型上实现42%的运动动态与物体生成准确率提升。
  • 适合需要快速部署多物体视频生成的开发者或创意工作者。

基于扩散模型的文本到视频生成近期取得显著进展,但在多物体视频生成方面仍面临挑战。现有模型常无法准确捕捉复杂物体交互,部分物体被误当作静态背景,且难以按提示生成多个独立物体,导致生成错误或特征混淆。本文提出一种无需训练的多物体视频生成新方法,利用扩散模型和大语言模型(LLM)的开放世界知识。我们以LLM作为物体轨迹的“导演”,通过噪声重初始化实现对真实运动的精确控制,并通过操纵注意力机制强化物体特异性特征与运动模式,防止跨物体特征干扰。大量实验验证了该方法在显著提升现有视频扩散模型多物体生成能力方面的有效性,相比基线模型,在运动动态和物体生成准确性上提升42%绝对值,同时保持高保真度与运动流畅性。

原文摘要 · Abstract (English)

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex object interactions, often treating some objects as static background elements and limiting their movement. In addition, they often fail to generate multiple distinct objects as specified in the prompt, resulting in incorrect generations or mixed features across objects. In this paper, we present a novel training-free approach for multi-object video generation that leverages the open world knowledge of diffusion models and large language models (LLMs). We use an LLM as the ``director'' of object trajectories, and apply the trajectories through noise re-initialization to achieve precise control of realistic movements. We further refine the generation process by manipulating the attention mechanism to better capture object-specific features and motion patterns, and prevent cross-object feature interference. Extensive experiments validate the effectiveness of our training free approach in significantly enhancing the multi-object generation capabilities of existing video diffusion models, resulting in 42% absolute improvement in motion dynamics and object generation accuracy, while also maintaining high fidelity and motion smoothness.

视频生成多物体扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。