构建15万张多图合成数据集,推动图像一致性生成研究
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
- 系统分类7类多图合成任务,构建高质量图文指令对
- 生成15万张一致性强的合成图像,含1.1万张真实拆解重构样本
- 提供新评测基准与加权参考评估指标,适合图像生成研究者
在可控图像生成中,从多个参考输入合成连贯一致的图像(即多图合成,MICo)仍具挑战性,主要受限于高质量训练数据的缺乏。为此,我们系统研究了MICo,将其分为7类代表性任务,并收集大规模高质量源图像,构建多样化MICo提示。利用强大的专有模型,合成大量均衡的复合图像,再经人机协同过滤与优化,形成包含身份一致性的MICo-150K数据集。我们进一步构建分解-重构(De&Re)子集,将11,000张真实复杂图像拆解为组件并重新组合,支持真实与合成双重生成。为实现全面评估,构建包含每任务100例、共300个高难度De&Re案例的MICo-Bench,并提出专用于MICo评估的新指标——加权参考VIEScore。最后,在MICo-150K上微调多个模型并在MICo-Bench上评估,结果表明该数据集有效赋能无MICo能力的模型,并提升已有能力模型的表现。值得注意的是,我们的基线模型Qwen-MICo(基于Qwen-Image-Edit微调)在三图合成上媲美Qwen-Image-2509,且支持任意多图输入,突破后者限制。本数据集、基准与基线共同为多图合成研究提供宝贵资源。
原文摘要 · Abstract (English)
In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo), remains a challenging problem, partly hindered by the lack of high-quality training data. To bridge this gap, we conduct a systematic study of MICo, categorizing it into 7 representative tasks and curate a large-scale collection of high-quality source images and construct diverse MICo prompts. Leveraging powerful proprietary models, we synthesize a rich amount of balanced composite images, followed by human-in-the-loop filtering and refinement, resulting in MICo-150K, a comprehensive dataset for MICo with identity consistency. We further build a Decomposition-and-Recomposition (De&Re) subset, where 11K real-world complex images are decomposed into components and recomposed, enabling both real and synthetic compositions. To enable comprehensive evaluation, we construct MICo-Bench with 100 cases per task and 300 challenging De&Re cases, and further introduce a new metric, Weighted-Ref-VIEScore, specifically tailored for MICo evaluation. Finally, we fine-tune multiple models on MICo-150K and evaluate them on MICo-Bench. The results show that MICo-150K effectively equips models without MICo capability and further enhances those with existing skills. Notably, our baseline model, Qwen-MICo, fine-tuned from Qwen-Image-Edit, matches Qwen-Image-2509 in 3-image composition while supporting arbitrary multi-image inputs beyond the latter's limitation. Our dataset, benchmark, and baseline collectively offer valuable resources for further research on Multi-Image Composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。