arXiv:2509.24285cs.AIcs.CL2025-09被引 5

构建跨学科科学问答验证基准与推理增强模型,提升大模型在科研场景的可信度。

SCI-Verifier: Scientific Verifier with Thinking

  • 设计跨学科验证基准SCI-VerifyBench,融合真实大模型输出与领域等价转换。
  • 提出SCI-Verifier模型,通过后训练实现逻辑推理与等价判断能力。
  • 适用于需高可靠性验证的科研大模型应用,如论文写作、试题评测。

随着大语言模型在科学推理中的广泛应用,答案格式复杂性和表达多样性使答案验证成为关键但具有挑战性的任务。现有科学领域验证研究存在两大局限:(a)缺乏系统评估标准且学科覆盖不足,难以全面评估;(b)严重依赖繁琐规则设计或提示工程,在复杂推理场景中效果不佳,限制跨学科泛化能力。为此,我们在数据与模型层面提出解决方案。数据层面,构建SCI-VerifyBench,一个涵盖数学、物理、生物、化学及通用科学问答的跨学科基准。该基准基于真实大模型输出,并通过领域特定等价变换生成具有挑战性且真实的测试数据。模型层面,强调推理对验证的重要性,提出SCI-Verifier——一种统一的推理增强型科学验证器。经后训练,SCI-Verifier展现出强大的逻辑推理与等价判断能力,同时保持输出简洁稳定。SCI-VerifyBench与SCI-Verifier共同构成科学验证的系统性框架,为科学领域大模型的可靠性与可应用性提供坚实支撑。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly applied to scientific reasoning, the complexity of answer formats and the diversity of equivalent expressions make answer verification a critical yet challenging task. Existing verification studies in scientific domains suffer from two major limitations: (a) the absence of systematic evaluation standards and insufficient disciplinary coverage, which hinders their comprehensive assessment; and (b) heavy reliance on cumbersome rule design or prompt engineering, which reduces their effectiveness in complex reasoning scenarios or limits their cross-disciplinary generalization. To address these challenges, we propose solutions at both the data and model levels. On the data side, we construct SCI-VerifyBench, a cross-disciplinary benchmark covering mathematics, physics, biology, chemistry, and general scientific QA. The benchmark is built from real LLM responses and enhanced with domain-specific equivalence transformations that generate challenging and realistic data. Model-based and expert annotations ensure both quality and diversity, enabling rigorous evaluation of verification ability. On the model side, we emphasize the importance of reasoning for verification and introduce SCI-Verifier, a unified reasoning-augmented verifier for scientific domains. Through post-training, SCI-Verifier demonstrates strong logical reasoning and equivalence judgment capabilities while maintaining concise and stable outputs. Together, SCI-VerifyBench and SCI-Verifier provide a principled framework for scientific verification, offering both systematic evaluation and practical pathways to enhance the reliability and applicability of LLMs in scientific domains.

科学验证大模型推理增强跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。