arXiv:2512.05119cs.IRcs.AI2025-12NeurIPS被引 4

为开放域问答中的图文混合生成设计新评测基准,解决现有评估不全面问题。

RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering

  • 构建基于检索增强生成的图文交错生成评测集,融合社交平台真实数据。
  • 提出文本与图像质量及一致性联合评估指标,与人工评分高度相关。
  • 支持开源与商用多模态大模型对比,验证训练有效性,适合研究图文生成者。

在真实场景中,视觉增强的回答能显著提升理解与记忆效果,凸显图文交错生成的价值。尽管已有进展,如将文本与图像处理统一于单一Transformer架构的视觉自回归模型,高质量图文交错内容生成仍具挑战。此外,现有评测多依赖单模态指标,难以准确评估图文融合输出的复杂性。为此,我们提出RAG-IGBench,一个专为开放域问答中基于检索增强生成(RAG-IG)的图文交错生成任务设计的综合性评测基准。RAG-IG结合多模态大语言模型(MLLMs)与检索机制,使模型可调用外部图文信息以生成连贯的多模态内容。不同于以往数据集,RAG-IGBench采用最新公开社交平台内容,并引入创新评估指标,综合衡量文本、图像质量及其一致性。通过对先进MLLMs(含开源与专有模型)在RAG-IGBench上的广泛实验,我们深入分析了模型能力与局限。同时,通过高相关性验证评估指标与人工评分的一致性。在RAG-IGBench训练集上微调的模型在多个基准上表现提升,证实本数据集的质量与实用性。基准代码已开源:https://github.com/USTC-StarTeam/RAG-IGBench。

原文摘要 · Abstract (English)

In real-world scenarios, providing user queries with visually enhanced responses can considerably benefit understanding and memory, underscoring the great value of interleaved image-text generation. Despite recent progress, like the visual autoregressive model that unifies text and image processing in a single transformer architecture, generating high-quality interleaved content remains challenging. Moreover, evaluations of these interleaved sequences largely remain underexplored, with existing benchmarks often limited by unimodal metrics that inadequately assess the intricacies of combined image-text outputs. To address these issues, we present RAG-IGBench, a thorough benchmark designed specifically to evaluate the task of Interleaved Generation based on Retrieval-Augmented Generation (RAG-IG) in open-domain question answering. RAG-IG integrates multimodal large language models (MLLMs) with retrieval mechanisms, enabling the models to access external image-text information for generating coherent multimodal content. Distinct from previous datasets, RAG-IGBench draws on the latest publicly available content from social platforms and introduces innovative evaluation metrics that measure the quality of text and images, as well as their consistency. Through extensive experiments with state-of-the-art MLLMs (both open-source and proprietary) on RAG-IGBench, we provide an in-depth analysis examining the capabilities and limitations of these models. Additionally, we validate our evaluation metrics by demonstrating their high correlation with human assessments. Models fine-tuned on RAG-IGBench's training set exhibit improved performance across multiple benchmarks, confirming both the quality and practical utility of our dataset. Our benchmark is available at https://github.com/USTC-StarTeam/RAG-IGBench.

图文生成评测基准RAG多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。