arXiv:2603.25319cs.CV2026-03被引 2

构建40万样本多参考图像生成数据集,解决参考图增多时模型性能下降问题。

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

  • 构建包含最多10张参考图的结构化长上下文数据集,覆盖四大任务维度。
  • 在4000样本基准上验证,多参考生成一致性显著提升。
  • 适合研究多参考图像生成、视觉推理与长上下文建模的学者使用。

多参考图像生成对多主体构图、叙事插画和新视角合成等真实场景至关重要,但现有模型在参考图数量增加时性能严重下降。我们发现根源在于数据瓶颈:现有数据集以单或少数参考图为主,缺乏结构化长上下文监督来学习密集的参考间依赖关系。为此,我们提出MacroData,一个包含40万样本的大规模数据集,每个样本最多包含10张参考图,系统地涵盖定制化、插画、空间推理和时间动态四个互补维度,全面覆盖多参考生成空间。针对评估标准缺失,我们进一步提出MacroBench,一个包含4000样本的基准,评估不同任务维度和输入规模下的生成一致性。大量实验表明,基于MacroData微调可显著提升多参考生成能力,消融实验揭示跨任务联合训练的协同优势及处理长上下文复杂性的有效策略。数据集与基准将公开发布。

原文摘要 · Abstract (English)

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance degradation as the number of input references grows. We identify the root cause as a fundamental data bottleneck: existing datasets are dominated by single- or few-reference pairs and lack the structured, long-context supervision needed to learn dense inter-reference dependencies. To address this, we introduce MacroData, a large-scale dataset of 400K samples, each containing up to 10 reference images, systematically organized across four complementary dimensions -- Customization, Illustration, Spatial reasoning, and Temporal dynamics -- to provide comprehensive coverage of the multi-reference generation space. Recognizing the concurrent absence of standardized evaluation protocols, we further propose MacroBench, a benchmark of 4,000 samples that assesses generative coherence across graded task dimensions and input scales. Extensive experiments show that fine-tuning on MacroData yields substantial improvements in multi-reference generation, and ablation studies further reveal synergistic benefits of cross-task co-training and effective strategies for handling long-context complexity. The dataset and benchmark will be publicly released.

多参考生成长上下文数据集图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。