arXiv:2502.04176cs.LGcs.IR2025-02被引 24

构建首个多模态生成评估基准,推动图文联合生成发展

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation

  • 提出MRAMG任务,实现图文混合答案生成
  • 构建含4800组问答的多领域多模态数据集
  • 支持11个主流模型评估,提供统计与LLM双指标

近期检索增强生成(RAG)进展显著提升了大语言模型的回答准确性和相关性,但现有方法仍以生成纯文本为主,即使在多模态检索增强生成(MRAG)场景中也仅生成文本答案。为此,我们提出多模态检索增强多模态生成(MRAMG)任务,旨在生成包含文本与图像的复合答案,充分挖掘语料中的多模态信息。针对该任务缺乏全面评估基准的问题,我们构建了MRAMG-Bench,一个精心策划、人工标注的基准,包含4,346篇文档、14,190张图像和4,800组问答对,涵盖网页、学术、生活三个领域,分为六个数据集,包含多样难度及复杂多图场景。为支持严谨评估,该基准提供统计与基于LLM的综合度量指标。此外,我们提出一种高效灵活的多模态生成框架,可调用LLM/MLLM生成多模态响应。所有数据集及11个生成模型的完整评估结果已公开于https://github.com/MRAMG-Bench/MRAMG。

原文摘要 · Abstract (English)

Recent advances in Retrieval-Augmented Generation (RAG) have significantly improved response accuracy and relevance by incorporating external knowledge into Large Language Models (LLMs). However, existing RAG methods primarily focus on generating text-only answers, even in Multimodal Retrieval-Augmented Generation (MRAG) scenarios, where multimodal elements are retrieved to assist in generating text answers. To address this, we introduce the Multimodal Retrieval-Augmented Multimodal Generation (MRAMG) task, in which we aim to generate multimodal answers that combine both text and images, fully leveraging the multimodal data within a corpus. Despite growing attention to this challenging task, a notable lack of a comprehensive benchmark persists for effectively evaluating its performance. To bridge this gap, we provide MRAMG-Bench, a meticulously curated, human-annotated benchmark comprising 4,346 documents, 14,190 images, and 4,800 QA pairs, distributed across six distinct datasets and spanning three domains: Web, Academia, and Lifestyle. The datasets incorporate diverse difficulty levels and complex multi-image scenarios, providing a robust foundation for evaluating the MRAMG task. To facilitate rigorous evaluation, MRAMG-Bench incorporates a comprehensive suite of both statistical and LLM-based metrics, enabling a thorough analysis of the performance of generative models in the MRAMG task. Additionally, we propose an efficient and flexible multimodal answer generation framework that can leverage LLMs/MLLMs to generate multimodal responses. Our datasets and complete evaluation results for 11 popular generative models are available at https://github.com/MRAMG-Bench/MRAMG.

多模态生成检索增强评估基准图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。