arXiv:2601.04651cs.AIcs.IR2026-01ACL被引 2

让推理模型互挑毛病,提升文档问答的深度与准确性。

Adversarial Yet Cooperative: Multi-Perspective Reasoning in Retrieved-Augmented Language Models

  • 设计推理-验证双角色对抗协作机制,促进多视角思考。
  • 在多个基准上显著优于基线模型,提升推理质量。
  • 无需外部评分模型,通过内部不确定性引导优化过程。

近期将大型推理模型(LRMs)与检索增强生成(RAG)结合的研究取得了显著进展,但仍面临两大挑战:(1)推理模型通常仅从单一、未受挑战的视角出发,难以对检索到的外部文档进行深入的自我修正推理;(2)现有训练范式过度依赖结果导向的奖励信号,无法有效引导复杂的多步推理过程。为此,我们提出一种名为对抗性推理RAG(ARR)的推理-验证框架。该框架中,推理器与验证器基于检索证据相互进行推理与批判,同时在过程感知优势的指导下运行,无需依赖外部评分模型。该奖励机制融合显式观测信号与模型内部不确定性,联合优化推理的准确性与验证的严格性。在多个基准测试上的实验表明,该方法有效提升了推理性能。

原文摘要 · Abstract (English)

Recent advances in synergizing large reasoning models (LRMs) with retrieval-augmented generation (RAG) have shown promising results, yet two critical challenges remain: (1) reasoning models typically operate from a single, unchallenged perspective, limiting their ability to conduct deep, self-correcting reasoning over external documents, and (2) existing training paradigms rely excessively on outcome-oriented rewards, which provide insufficient signal for shaping the complex, multi-step reasoning process. To address these issues, we propose an Reasoner-Verifier framework named Adversarial Reasoning RAG (ARR). The Reasoner and Verifier engage in reasoning on retrieved evidence and critiquing each other's logic while being guided by process-aware advantage that requires no external scoring model. This reward combines explicit observational signals with internal model uncertainty to jointly optimize reasoning fidelity and verification rigor. Experiments on multiple benchmarks demonstrate the effectiveness of our method.

推理增强对抗训练多视角推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。