arXiv:2604.18176cs.AIquant-ph2026-04ACL被引 1

用物理一致数据与验证感知强化学习提升大模型科学推理能力

QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning

论文配图:QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning
图 1 · 摘自论文原文
  • 构建量子物理一致性数据集,结合确定性求解与语义审计保证严谨性
  • 提出验证感知奖励模型,动态融合执行结果与语义评估实现精准监督
  • 8B模型性能媲美专有模型,证明规则反馈可高效替代纯规模扩展

大语言模型在通用推理中表现强劲,但在量子力学等科学领域常因缺乏可验证训练资源及粗粒度反馈而不可靠。为此,我们提出QuantumQA,一个通过任务自适应策略和混合验证协议构建的大规模数据集,结合确定性求解器与语义审计确保科学严谨性。在此基础上,我们设计验证感知奖励模型(VRM),用于可验证奖励的强化学习(RLVR),采用自适应奖励融合(ARF)机制,动态整合科学执行套件(SES)的确定性信号与多维语义评估,实现精确监督。实验表明,该方法持续优于基线与通用偏好模型。特别地,优化后的8B模型性能媲美专有模型,验证了将可验证、规则化反馈引入强化学习循环,是一种高效的参数节省型替代方案。

原文摘要 · Abstract (English)

Large language models (LLMs) show strong capabilities in general reasoning but typically lack reliability in scientific domains like quantum mechanics, which demand strict adherence to physical constraints. This limitation arises from the scarcity of verifiable training resources and the inadequacy of coarse feedback signals in standard alignment paradigms. To address the data challenge, we introduce QuantumQA, a large-scale dataset constructed via a task-adaptive strategy and a hybrid verification protocol that combines deterministic solvers with semantic auditing to guarantee scientific rigor. Building on this foundation, we propose the verification-aware reward model (VRM) tailored for Reinforcement Learning with Verifiable Rewards (RLVR), which employs an adaptive reward fusion (ARF) mechanism to dynamically integrate deterministic signals from a scientific execution suite (SES) with multidimensional semantic evaluations for precise supervision. Experimental results demonstrate that our method consistently outperforms baselines and general-purpose preference models. Notably, our optimized 8B model achieves performance competitive with proprietary models, validating that incorporating verifiable, rule-based feedback into the reinforcement learning loop offers a parameter-efficient alternative to pure scaling.

科学推理强化学习量子物理数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。