新基准多参考图像生成,测试模型处理8张不同参考图的能力。
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- 设计多参考图像生成评估基准,支持最多8张参考图
- 覆盖跨域、尺度、罕见概念等5类挑战性场景
- 适合研究多参考图像生成的模型性能与缺陷
近期文本到图像生成模型已具备多参考生成与编辑能力,即从多张参考图中继承主体外观并在新场景中重绘。然而现有基准数据集通常仅关注单张或少数几张参考图,难以衡量模型在大量参考下的表现或识别其弱点。此外,任务定义模糊,局限于编辑内容或参考数量等简单维度,无法捕捉异构参考融合的复杂挑战。为此,我们提出 MultiBanana 基准,全面覆盖多参考设置中的五大核心问题:(1)参考图数量变化(最多8张),(2)参考图间领域不匹配(如照片与动漫),(3)参考与目标场景尺度差异,(4)参考图含罕见概念(如红色香蕉),(5)多语言文本参考。对多种文本到图像模型的分析揭示了其性能差异、典型失败模式及改进方向。MultiBanana 已开源,旨在推动该领域发展并建立公平比较标准。数据与代码见 https://github.com/matsuolab/multibanana。
原文摘要 · Abstract (English)
Recent text-to-image generation models have acquired the ability of multi-reference generation and editing; that is, to inherit the appearance of subjects from multiple reference images and re-render them in new contexts. However, existing benchmark datasets often focus on generation using a single or a few reference images, which prevents us from measuring progress in model performance or identifying weaknesses when following instructions with a larger number of references. In addition, their task definitions are still vague, limited to axes such as ``what to edit'' or ``how many references are given'', and therefore fail to capture the challenges inherent in combining heterogeneous references. To address this gap, we introduce MultiBanana, which is designed to assess the edge of model capabilities by widely covering problems specific to multi-reference settings: (1) varying the number of references (up to 8), (2) domain mismatch among references (e.g., photo vs. anime), (3) scale mismatch between reference and target scenes, (4) references containing rare concepts (e.g., a red banana), and (5) multilingual textual references for rendering. Our analysis among a variety of text-to-image models reveals their respective performances, typical failure modes, and areas for improvement. MultiBanana is released as an open benchmark to push the boundaries and establish a standardized basis for fair comparison in multi-reference image generation. Our data and code are available at https://github.com/matsuolab/multibanana .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。