arXiv:2604.24564cs.CLcs.IR2026-04被引 1

提出新指标MEG,让多模态检索更精准判断证据是否真有用。

MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG

论文配图:MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG
图 1 · 摘自论文原文
  • 用语义核心锚定法识别关键信息词,量化证据真实贡献度。
  • 在M²RAG上超越基线,生成答案更准确且多模态一致。
  • 适合研究知识增强生成、多模态推理的开发者与学者。

多模态检索增强生成(MRAG)解决了多模态大模型幻觉和知识过时问题。但现有系统难以区分检索到的多模态数据是否真正支撑答案核心,还是仅表面相关。现有评估指标依赖启发式位置置信度,无法捕捉多模态实体的信息密度。为此,我们提出多模态证据接地(MEG),一种语义感知指标,用于量化检索证据的贡献。不同于标准置信度,MEG采用语义确定性锚定,聚焦高逆文档频率(IDF)的信息承载词,更好捕捉答案语义核心。基于MEG,我们构建MEG-RAG框架,训练多模态重排序器以对齐检索证据与真实答案的语义锚点。通过优先选择基于语义接地的高价值内容而非词概率分布,MEG-RAG提升了生成输出的准确性和多模态一致性。在M²RAG基准上的大量实验表明,MEG-RAG持续优于强基线,并在不同教师模型间展现出鲁棒泛化能力。

原文摘要 · Abstract (English)

Multimodal Retrieval-Augmented Generation (MRAG) addresses key limitations of Multimodal Large Language Models (MLLMs), such as hallucination and outdated knowledge. However, current MRAG systems struggle to distinguish whether retrieved multimodal data truly supports the semantic core of an answer or merely provides superficial relevance. Existing metrics often rely on heuristic position-based confidence, which fails to capture the informational density of multimodal entities. To address this, we propose Multi-modal Evidence Grounding (MEG), a semantic-aware metric that quantifies the contribution of retrieved evidence. Unlike standard confidence measures, MEG utilizes Semantic Certainty Anchoring, focusing on high-IDF information-bearing tokens that better capture the semantic core of the answer. Building on MEG, we introduce MEG-RAG, a framework that trains a multimodal reranker to align retrieved evidence with the semantic anchors of the ground truth. By prioritizing high-value content based on semantic grounding rather than token probability distributions, MEG-RAG improves the accuracy and multimodal consistency of generated outputs. Extensive experiments on the M$^2$RAG benchmark show that MEG-RAG consistently outperforms strong baselines and demonstrates robust generalization across different teacher models.

多模态RAG证据选择语义评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。