用强化学习让模型主动发现论文错误并给出证据。
Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper

- 设计三步验证流程:思考-验证-扫描,驱动模型主动检查
- 构建1.29万样本数据集,覆盖6类科学错误,支持链式推理
- 奖励机制细化到推理完整性和证据匹配度,适合科研助手开发
多模态大语言模型日益成为科研助手,但仍难以实现完全自主。要实现这一目标,模型需能主动查阅学术论文,构建全局证据视图,并在无预设问题或证据的情况下做出可追溯的判断。然而,现有研究在无具体问题与证据的验证任务范式和训练方法上仍十分有限。本文通过科学错误检测这一任务来探索该挑战,要求模型判断是否存在错误,并以证据为基础进行推理。为此,我们提出 VERA-RL,一种基于强化学习的科学错误检测框架。遵循‘思考—验证—扫描’的流程,我们构建了包含12,900个样本的VERA-13K数据集,按4,300组匹配链组织,涵盖研究全流程中的6类科学错误,覆盖广泛的自然科学领域。我们还引入细粒度奖励机制,分别评估推理完整性、证据对齐度和错误精确性。使用Qwen3-VL-8B在VERA-RL上训练后,模型的可验证推理能力显著提升,接近旗舰多模态模型Gemini 3 Pro和Qwen3-VL-235B-A22B的扫描表现。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly capable scientific assistants, yet they remain far from fully autonomous research. This transition requires models to actively inspect academic papers, build global evidence views, and make traceable judgments without prespecified issues or evidence. However, existing work provides limited task paradigms or training studies for such issue- and evidence-absent verification. We study this challenge through scientific error detection, where models must determine whether errors exist and justify them with evidence-based reasoning. To fill this gap, we present VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers. Following a Reason--Verify--Scan progression, we construct VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains. We further introduce fine-grained rewards for reasoning completeness, evidence alignment, and error precision. Training Qwen3-VL-8B with VERA-RL substantially improves verifiable reasoning, approaching flagship MLLMs such as Gemini 3 Pro and Qwen3-VL-235B-A22B on Scan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。