arXiv:2512.17194cs.AI2025-12AAAI

用两阶段强化学习让多模态生成更可解释

MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation

  • 分两阶段用强化学习优化图文检索与回答生成
  • 在WebQA和MultimodalQA上达领先效果
  • 适合关注模型可解释性的研究者

多模态检索增强生成(MMRAG)通过整合外部多模态知识,实现高可信度生成,在复杂多模态场景中表现优异。然而现有方法难以揭示检索与生成背后的推理逻辑,限制了结果的可解释性。为此,我们引入强化学习,提出两阶段强化微调框架,提升多模态大语言模型的推理能力,实现可解释的多模态检索增强生成。第一阶段采用基于规则的强化微调,对多模态文档进行粗粒度逐项排序,有效过滤明显无关内容;第二阶段采用基于推理的强化微调,联合优化细粒度列表级排序与答案生成,引导模型输出可解释的推理过程。该方法在WebQA和MultimodalQA两个基准数据集上取得当前最优性能,且通过全面消融实验验证了有效性。

原文摘要 · Abstract (English)

Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments.

多模态生成强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。