arXiv:2411.17188cs.CVcs.CL2024-11中稿 · ICLR被引 19

提出多粒度评估框架ISG,解决图文混排生成的一致性难题

Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment

  • 用场景图结构分析文本与图像块间关系,分四层评估
  • 统一模型在整体层面表现差,组合模型提升111%但局部仍弱
  • 提供带黄金答案的基准数据集,适合视觉任务研究者使用

许多真实用户查询(如“如何做蛋炒饭?”)需要同时生成文字步骤和配套图片的系统支持,类似食谱。当前能生成图文混排内容的模型面临模态内与跨模态一致性挑战。为此,我们提出ISG——一种针对图文混排生成的综合性评估框架。ISG利用场景图结构捕捉文本与图像块之间的关系,在四个粒度层次(整体、结构、块级、图像特定)进行评估,实现对一致性、连贯性和准确性的精细判断,并提供可解释的问答反馈。同时,我们构建了涵盖8个类别、21个子类别的基准数据集ISG-Bench,共包含1,150个样本,包含复杂的语言-视觉依赖关系及黄金答案,用于有效评估视觉主导任务(如风格迁移)。实验表明,现有统一视觉-语言模型在生成混排内容上表现不佳;而采用分离语言与图像模型的组合方法,在整体层面相比统一模型提升111%,但在块级与图像级仍不理想。为推动后续研究,我们开发了基于“规划-执行-优化”流程的ISG-Agent基线代理,通过调用工具实现122%的性能提升。

原文摘要 · Abstract (English)

Many real-world user queries (e.g. "How do to make egg fried rice?") could benefit from systems capable of generating responses with both textual steps with accompanying images, similar to a cookbook. Models designed to generate interleaved text and images face challenges in ensuring consistency within and across these modalities. To address these challenges, we present ISG, a comprehensive evaluation framework for interleaved text-and-image generation. ISG leverages a scene graph structure to capture relationships between text and image blocks, evaluating responses on four levels of granularity: holistic, structural, block-level, and image-specific. This multi-tiered evaluation allows for a nuanced assessment of consistency, coherence, and accuracy, and provides interpretable question-answer feedback. In conjunction with ISG, we introduce a benchmark, ISG-Bench, encompassing 1,150 samples across 8 categories and 21 subcategories. This benchmark dataset includes complex language-vision dependencies and golden answers to evaluate models effectively on vision-centric tasks such as style transfer, a challenging area for current models. Using ISG-Bench, we demonstrate that recent unified vision-language models perform poorly on generating interleaved content. While compositional approaches that combine separate language and image models show a 111% improvement over unified models at the holistic level, their performance remains suboptimal at both block and image levels. To facilitate future work, we develop ISG-Agent, a baseline agent employing a "plan-execute-refine" pipeline to invoke tools, achieving a 122% performance improvement.

图文生成评估框架场景图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。