通过强化学习提升多模态生成中的引用准确性和答案质量。
MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning

- 引入动态迭代检索与裁剪,让引用成为持续优化的推理过程。
- 结合过程与结果反馈,在多个基准上实现引用精度与答案质量同步提升。
- 适合关注多模态可追溯生成、可信AI的开发者与研究者。
多模态检索增强生成(RAG)中引入视觉引用对确保多模态大模型(MLLMs)输出的可追溯性与可验证性至关重要。然而,现有基于RAG和监督微调(SFT)的方法难以实现鲁棒的跨模态推理,导致视觉引用不准确或引用与答案脱节。为此,我们提出MCite-RL,一种增强引用的代理式强化学习框架,用于可靠的多模态RAG。MCite-RL引入代理精炼模块,通过迭代检索、推理与递归裁剪逐步缩小搜索空间,将引用转化为动态、证据驱动的推理过程,而非静态步骤。此外,我们设计了引用增强型奖励机制,在强化学习框架内整合过程级与结果级反馈,联合优化答案准确率与来源可追溯性。在Wiki-VISA、FinRAGBench-V和MMLongBench-Doc等多个基准上的实验表明,MCite-RL能有效实现引用精确度与答案质量的联合优化。
原文摘要 · Abstract (English)
Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning, causing imprecise visual citations or decoupling between the citation and the generated answers. To address these limitations, we propose MCite-RL, a citation-enhanced agentic reinforcement learning framework designed for reliable multimodal RAG. MCite-RL introduces an Agentic Refinement module for visual citation that employs iterative retrieval, reasoning, and recursive cropping to progressively narrow the search space, transforming citation into a dynamic, evidence-driven reasoning process rather than a static step. Furthermore, we incorporate a Citation-enhanced Reward mechanism that integrates both process-level and outcome-level feedback within a reinforcement learning paradigm to jointly optimize answer accuracy and source traceability. Extensive experiments on benchmarks such as Wiki-VISA, FinRAGBench-V, and MMLongBench-Doc demonstrate that MCite-RL effectively achieves the joint optimization of citation precision and answer quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。