构建可验证的科学发现环境,让大模型在真实数据上高效试错。
Scaling Scientific Discovery Environments for Turn-Level Agentic RL

- 用可验证流程构建科学分析环境,支持多轮交互。
- 在基准测试中,140亿参数模型达到当前最优水平。
- 适合需要严谨推理的科研自动化研究者使用。
大型语言模型代理在数据驱动的科学发现任务中展现出潜力,代理与执行环境互动并生成统计结论。然而,长周期科学分析仍受限于缺乏对真实科学数据的流程监督环境。本文提出SciDisco,一个可在真实科学数据上训练可验证科学发现代理的可扩展框架。SciThèque将假设、数据集、隐藏证据图谱和验证器整合为任务环境,使分析进展能在交互过程中被检查。基于有向无环图(DAG)的轨迹合成利用这些环境生成经验证器过滤的多轮示范。DiscoPO则以环境为训练信号源,为产生可验证分析证据的动作分配回合级奖励。实验表明,SciDisco-14B在假设驱动的科学数据分析基准上达到当前最佳性能。
原文摘要 · Abstract (English)
Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciThèque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。