arXiv:2510.24870cs.CLcs.CV2025-10被引 1

为多模态检索生成设计评估框架,解决现有方法只看文字的局限。

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

  • 提出基于事实和引用的双指标评估框架,适配音视频等多模态信息。
  • 人工评估显示该框架与输出质量判断高度一致,验证有效性。
  • 开源工具链支持自动评估,推动多模态RAG研究标准化。

我们提出MiRAGE,一种面向多模态来源的检索增强生成(RAG)评估框架。随着音视频内容成为网络信息的主要来源,RAG系统需整合此类多模态信息进行生成,但现有评估仍以文本为中心,难以适用。MiRAGE采用以论断为中心的评估思路,包含InfoF1(评估事实性与信息覆盖度)和CiteF1(评估引用支持与完整性)。人工评估表明,该框架与输出质量的外部判断高度一致。我们还提供了自动实现版本,并构建了三种主流文本RAG指标(ALCE、ARGUE、RAGAS)的多模态变体,揭示了文本中心评估的局限性,为自动化评估奠定基础。相关代码与评估方法已开源。

原文摘要 · Abstract (English)

We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a prevalent source of information online, it is essential for RAG systems to integrate information from these sources into generation. However, existing evaluations for RAG are text-centric, limiting their applicability to multimodal settings. MiRAGE is a claim-centric approach to multimodal RAG evaluation, consisting of InfoF1, which assesses factuality and information coverage, and CiteF1, which assesses citation support and completeness. We show that, when applied by humans, MiRAGE strongly aligns with extrinsic judgments of output quality. We additionally introduce an automatic implementation of MiRAGE as well as multimodal variants of three prominent text-based RAG metrics -- ALCE, ARGUE, and RAGAS -- demonstrating the limitations of text-centric work and laying the groundwork for automatic evaluation. We release open-source implementations and outline evaluation methods for multimodal RAG.

多模态RAG评估信息抽取自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。