arXiv:2508.06905cs.CV2025-08中稿 · ACM MM 2025 Datase…被引 8

让AI同时参考多张图生成新图像,更接近人类创作方式。

MultiRef: Controllable Image Generation with Multiple Visual References

  • 设计时融合多张参考图,突破单图输入限制。
  • 顶尖模型在合成数据上仅达66.6%准确率,真实场景下79.0%。
  • 提供3.8万张高质量图像数据集,适合创意生成研究者使用。

视觉设计师常从多个视觉参考中汲取灵感,融合不同元素与美学原则创作作品。然而,当前图像生成框架大多依赖单一输入——要么是文本提示,要么是单张参考图。本文聚焦于利用多视觉参考进行可控图像生成的任务。我们提出了MultiRef-bench评估框架,包含990个合成样本和1000个真实世界样本,要求模型整合多张参考图的视觉内容。合成样本由我们的数据引擎RefBlend生成,涵盖10种参考类型和33种组合方式。基于RefBlend,我们进一步构建了包含3.8万张高质量图像的MultiRef数据集,以推动后续研究。在三种交叉图像-文本模型(OmniGen、ACE、Show-o)和六种代理框架(如ChatDiT和LLM + SD)上的实验表明,即使最先进的系统在多参考条件下的表现仍有限:在合成样本上平均准确率为66.6%,真实案例中为79.0%(相比理想答案)。这些发现为开发更灵活、类人化的创作工具提供了重要方向。数据集已公开:https://multiref.github.io/。

原文摘要 · Abstract (English)

Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs -- either text prompts or individual reference images. In this paper, we focus on the task of controllable image generation using multiple visual references. We introduce MultiRef-bench, a rigorous evaluation framework comprising 990 synthetic and 1,000 real-world samples that require incorporating visual content from multiple reference images. The synthetic samples are synthetically generated through our data engine RefBlend, with 10 reference types and 33 reference combinations. Based on RefBlend, we further construct a dataset MultiRef containing 38k high-quality images to facilitate further research. Our experiments across three interleaved image-text models (i.e., OmniGen, ACE, and Show-o) and six agentic frameworks (e.g., ChatDiT and LLM + SD) reveal that even state-of-the-art systems struggle with multi-reference conditioning, with the best model OmniGen achieving only 66.6% in synthetic samples and 79.0% in real-world cases on average compared to the golden answer. These findings provide valuable directions for developing more flexible and human-like creative tools that can effectively integrate multiple sources of visual inspiration. The dataset is publicly available at: https://multiref.github.io/.

图像生成多参考创意设计数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。