Kaleido可基于多张参考图生成连贯视频,解决多人物一致性难题。
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
- 构建专用数据管道并引入R-RoPE编码,提升多参考图融合稳定性
- 在多个基准上超越已有方法,视频一致性与参考图像保真度显著提升
- 适合需要多角色视频生成的创作者或研究者使用
我们提出Kaleido,一种主体到视频(S2V)生成框架,旨在根据目标主体的多张参考图像合成保持主体一致性的视频。尽管近期S2V模型取得进展,现有方法在多主体一致性维护和背景解耦方面仍不足,常导致参考图像保真度下降和语义漂移。这主要源于训练数据缺乏多样性与高质量样本,以及跨配对数据(即来自不同实例的成对样本)缺失。此外,当前多参考图像整合机制不佳,易造成主体混淆。为此,我们设计专用数据构建流程,包含低质量样本过滤与多样数据合成,以生成保持一致性的训练数据;同时引入参考旋转位置编码(R-RoPE),实现稳定精确的多图像整合。大量实验表明,Kaleido在多个基准测试中显著优于先前方法,在一致性、保真度与泛化能力方面均有提升,推动了S2V生成技术的发展。
原文摘要 · Abstract (English)
We present Kaleido, a subject-to-video~(S2V) generation framework, which aims to synthesize subject-consistent videos conditioned on multiple reference images of target subjects. Despite recent progress in S2V generation models, existing approaches remain inadequate at maintaining multi-subject consistency and at handling background disentanglement, often resulting in lower reference fidelity and semantic drift under multi-image conditioning. These shortcomings can be attributed to several factors. Primarily, the training dataset suffers from a lack of diversity and high-quality samples, as well as cross-paired data, i.e., paired samples whose components originate from different instances. In addition, the current mechanism for integrating multiple reference images is suboptimal, potentially resulting in the confusion of multiple subjects. To overcome these limitations, we propose a dedicated data construction pipeline, incorporating low-quality sample filtering and diverse data synthesis, to produce consistency-preserving training data. Moreover, we introduce Reference Rotary Positional Encoding (R-RoPE) to process reference images, enabling stable and precise multi-image integration. Extensive experiments across numerous benchmarks demonstrate that Kaleido significantly outperforms previous methods in consistency, fidelity, and generalization, marking an advance in S2V generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。