arXiv:2608.22819cs.CV2026-08

对比三种无需训练的多主体图像转视频方法,揭示各自优劣。

Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation

论文配图:Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation
图 1 · 摘自论文原文
  • 分直接、并行、串行三类生成方式,解耦主体视觉与动作条件。
  • 并行法保形好但主体间互动弱,串行法上下文连贯但受顺序影响。
  • 实证分析不同场景下表现,为可控视频生成提供设计参考。

文本条件的图像转视频(I2V)生成进展迅速,但多主体视频生成仍具挑战:模型需同时保持各主体外观、分配不同运动,并维持空间与时间上的连贯交互。本文系统比较了三种无需训练的多主体I2V生成范式:直接、并行与串行生成。直接生成将完整参考图与提示输入预训练I2V模型,联合合成所有主体与动作;并行生成将图像与提示分解为个体主体的视觉与文本条件,独立生成各主体视频后拼接,降低每步复杂度但削弱主体间上下文;串行生成先生成背景视频,再逐步引入个体主体,保留累积场景上下文但对主体顺序敏感且易产生误差传播。我们在多样多主体场景中评估三类方法在外观保持、动作保真、时间一致性及主体间连贯性方面的表现,并分析其典型失败模式。结果揭示各类范式的优劣,为可控多主体视频生成系统的设计提供实践指导。

原文摘要 · Abstract (English)

Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must simultaneously preserve the appearance of each subject, assign distinct motions, and maintain coherent spatial and temporal interactions. This paper presents a systematic study of three representative paradigms for training-free multi-subject I2V generation: direct, parallel, and sequential generation. Direct generation applies a pretrained I2V model to the complete reference image and prompt, requiring all subjects and motions to be synthesized jointly. Parallel and sequential generation instead decompose the reference image and prompt into subject-specific visual and textual conditions. Parallel generation synthesizes each subject independently and subsequently composes the resulting videos, reducing the complexity of each generation step at the cost of weaker inter-subject context. Sequential generation first synthesizes a background video and then progressively introduces individual subjects. This preserves accumulated scene context but introduces sensitivity to subject ordering and error propagation. We empirically evaluate the three paradigms across diverse multi-subject scenes, comparing appearance preservation, motion fidelity, temporal consistency, and inter-subject coherence, while also characterizing their distinct failure modes. Our findings reveal the strengths and limitations of each paradigm and offer practical insights for designing controllable multi-subject video generation systems.

图像转视频多主体生成无训练生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。