arXiv:2608.30109cs.CL2026-08中稿 · EMNLP

让大模型学会模拟科学家思考过程,提升科研辅助能力

COGTRL: Training LLMs for Scientific Discovery Assistance using Cognitive Traces via Reinforcement Learning

论文配图:COGTRL: Training LLMs for Scientific Discovery Assistance using Cognitive Traces via Reinforcement Learning
图 1 · 摘自论文原文
  • 用强化学习联合优化思维过程和科研步骤,模拟真实研究推理
  • 在两个领域中方法质量平均提升7.85分,接近70B模型表现
  • 适合需要可解释科研建议的科研人员使用

大规模语言模型(LLMs)在科学发现辅助中日益重要,但多数论文忽略了研究过程中对约束条件、失败方案和迭代决策等细粒度认知过程的记录。这些认知过程对现实科学家在约束条件下达成目标至关重要。本文提出COGTRL,一种轨迹级强化学习框架,通过联合优化认知痕迹与科学步骤,在交错过程中模拟具认知基础的推理。在两个30亿参数模型及两个科学领域(人工智能与材料科学)中,COGTRL相比同类30亿参数基线平均提升方法质量7.85分,性能媲美700亿参数模型。领域专家评估显示,更偏好COGTRL生成的方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) trained on extensive scientific research are increasingly integrated as assistants for scientific discovery. However, most research papers omit the fine-grained cognitive process of examining constraints, failed alternatives, and iterative decisions required to achieve the desired goal. Such cognitive processes are vital for real-world scientists working toward specific goals under constraints. In this paper, we show that LLMs, when trained to produce such cognitive traces, perform better as scientific discovery assistants than when trained solely on scientific literature. We propose COGTRL, a trajectory-level reinforcement learning framework that trains LLMs to emulate cognitively grounded reasoning by jointly optimizing cognitive traces and the scientific steps produced in an interleaved manner. Across two 3B-parameter models and two scientific domains (AI and Materials Science), COGTRL improves method quality by an average of 7.85 points over comparable 3B model baselines and achieves competitive performance relative to 70B parameter models. Moreover, analysis by domain experts shows a preference for methods generated by COGTRL over the baselines.

大模型科研辅助强化学习认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。