arXiv:2601.17756cs.CVcs.AI2026-01International Conf…被引 8

多视角参考生成更一致的视频,突破单视图限制。

MV-S2V: Multi-View Subject-Consistent Video Generation

  • 用多视角参考提升视频主体3D一致性
  • 合成数据+真实数据训练,解决数据稀缺问题
  • 创新时移RoPE区分不同视角与主体,适合可控视频生成研究

现有主体到视频生成(S2V)方法虽能实现高保真与主体一致性,但仅限于单视角参考,实质退化为S2I+I2V流程,未能发挥视频主体控制潜力。本文提出并解决具有挑战性的多视角S2V(MV-S2V)任务,通过多参考视角合成视频以实现3D级主体一致性。针对训练数据稀缺,我们设计了一套定制化合成数据生成流程,并结合小规模真实采集数据增强训练效果。另一关键难点是条件生成中跨主体与跨视角参考的混淆问题,为此我们引入时移旋转位置编码(TS-RoPE),有效区分同一主体的不同视角与不同主体。所提框架在多视角参考下实现优越的3D主体一致性与高质量视觉输出,为可控视频生成开辟新方向。代码与数据已公开。

原文摘要 · Abstract (English)

Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to single-view subject references. This limitation renders the S2V task reducible to an S2I + I2V pipeline, failing to exploit the full potential of video subject control. In this work, we propose and address the challenging Multi-View S2V (MV-S2V) task, which synthesizes videos from multiple reference views to enforce 3D-level subject consistency. Regarding the scarcity of training data, we first develop a synthetic data curation pipeline to generate highly customized synthetic data, complemented by a small-scale real-world captured dataset to boost the training of MV-S2V. Another key issue lies in the potential confusion between cross-subject and cross-view references in conditional generation. To overcome this, we further introduce Temporally Shifted RoPE (TS-RoPE) to distinguish between different subjects and distinct views of the same subject in reference conditioning. Our framework achieves superior 3D subject consistency w.r.t. multi-view reference images and high-quality visual outputs, establishing a new meaningful direction for subject-driven video generation. Code and data are available at: https://szy-young.github.io/mv-s2v

视频生成多视角主体一致扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。