arXiv:2508.07313cs.CV2025-08AAAI被引 14

用强化学习提升多页文档理解,先找关键页再答题

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

  • 通过证据页引导的强化学习框架,分步推理更精准
  • 在4.8千样本训练集上达成顶尖多页理解性能
  • 适合需要跨页分析的科研与文档处理场景

多页文档理解对多模态大语言模型(MLLM)构成重大挑战,需细粒度视觉理解与跨页多跳推理。现有研究虽探索过强化学习(RL)提升模型高级推理能力,但其在多页文档理解中的应用仍不充分。本文提出DocR1,一种基于新型强化学习框架EviGRPO(Evidence Page-Guided GRPO)训练的MLLM。EviGRPO引入证据感知奖励机制,促进从粗到精的推理策略,引导模型先检索相关页面再生成答案。该训练范式使我们在有限监督下构建高质量模型。为此,我们设计两阶段标注流程与课程学习策略,构建两个数据集:包含4.8k样本的高质量训练集EviBench,以及基于学术论文的8.6k问答对评估基准ArxivFullQA。广泛实验表明,DocR1在多页任务上达到当前最优表现,同时在单页基准上保持强劲性能。

原文摘要 · Abstract (English)

Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reasoning across pages. While prior work has explored reinforcement learning (RL) for enhancing advanced reasoning in MLLMs, its application to multi-page document understanding remains underexplored. In this paper, we introduce DocR1, an MLLM trained with a novel RL framework, Evidence Page-Guided GRPO (EviGRPO). EviGRPO incorporates an evidence-aware reward mechanism that promotes a coarse-to-fine reasoning strategy, guiding the model to first retrieve relevant pages before generating answers. This training paradigm enables us to build high-quality models with limited supervision. To support this, we design a two-stage annotation pipeline and a curriculum learning strategy, based on which we construct two datasets: EviBench, a high-quality training set with 4.8k examples, and ArxivFullQA, an evaluation benchmark with 8.6k QA pairs based on scientific papers. Extensive experiments across a wide range of benchmarks demonstrate that DocR1 achieves state-of-the-art performance on multi-page tasks, while consistently maintaining strong results on single-page benchmarks.

文档理解强化学习多页推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。