arXiv:2607.15619cs.CV2026-07

用结构化信息提升多参考图像生成的准确性

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling

论文配图:StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
图 1 · 摘自论文原文
  • 采用字典式结构编码多参考图,明确指定生成意图
  • 在复杂指令下,语义对齐与细节一致性显著提升
  • 适用于需要精确控制生成内容的研究者与开发者

多参考图像生成旨在根据文本指令整合多张参考图的属性生成新图像。随着参考图数量增加,任务需理解复杂的语义关系,如正确关联属性与目标主体,并规划主体与环境间的合理空间布局。现有方法仅依赖自然语言指令,常因指令冗长模糊及高质量数据稀缺,导致语义错位与生成不一致。本文提出StructGen,采用类似字典的结构化格式编码多参考图像,实现生成意图的显式、无歧义表达。为此,我们基于真实高质量图像构建了结构化数据集,设计了相应训练框架及针对复杂多参考场景的基准测试。在公开基准和自建基准上的大量实验表明,StructGen在语义对齐和参考生成一致性上持续优于现有方法,尤其在含多个参考的复杂指令下表现突出。

原文摘要 · Abstract (English)

Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increases, the task necessitates complex semantic comprehension, such as correctly associating attributes with the intended subjects and planing out coherent spatial arrangement between subjects and their environments. Existing approaches, which rely solely on natural language instruction, often fail to capture these complex intentions precisely, leading to semantic misalignment and inconsistent generation. We identify two key factors behind these limitations: natural language instructions are often verbose and ambiguous, and high-quality multi-reference data is scarce. To address these issues, we propose StructGen, which employs a structured, dictionary-like format to encode multiple reference images, thereby enabling explicit and unambiguous specification of generation intentions. To support this design, we construct a structured dataset based on high-quality real images and develop a corresponding training framework, along with a dedicated benchmark for challenging multi-reference scenarios. Extensive experiments on both public benchmarks and our proposed benchmark demonstrate that StructGen consistently outperforms existing methods on both semantic alignment and detailed reference-generation consistency, especially under complex instructions with multiple references. The code is available at https://jianingpeng0382.github.io/StructGen/

图像生成多参考结构化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。