用结构化奖励机制让AI生成可执行的生物实验协议。
Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism
- 分步构建协议:先分析再结构化,最后表达,确保每步清晰可验证。
- 在27个生物领域12000+协议上训练,生成结果逻辑更顺、语义更准。
- 适合科研人员快速生成可复现实验流程,提升自动化研究效率。
可重复科学的基础在于精确、逻辑有序且可执行的实验协议。通过自然语言查询自主生成协议,能极大提升复现效率。然而当前主流大模型常生成不完整或不一致的协议,限制其应用。为此,我们首先构建了涵盖27个生物子领域的超12000条结构化协议数据集SciRecipe,包含理解与求解任务。提出“草图-填充”范式,将分析、结构化与表达分离,确保每一步明确可验证。同时设计结构化组件奖励机制,评估步骤粒度、操作顺序与语义一致性,使模型优化契合实验可靠性。基于此,我们开发了Thoth模型,通过知识到行动的阶段性训练,实现从知识获取到操作推理再到稳健可执行协议生成。在多个基准测试中,Thoth显著优于专有与开源大模型,在步骤对齐、逻辑排序与语义准确率上均有提升。该方法为连接知识与实验执行的可靠科学助手奠定基础。所有数据、代码与模型将公开发布。
原文摘要 · Abstract (English)
The foundation of reproducible science lies in protocols that are precise, logically ordered, and executable. The autonomous generation of these protocols through natural language queries could greatly improve the efficiency of the reproduction process. However, current leading large language models (LLMs) often generate incomplete or inconsistent protocols, limiting their utility. To address this limitation, we first introduce SciRecipe, a large-scale dataset of over 12K structured protocols spanning 27 biological subfields and encompassing both comprehension and problem-solving tasks. To further improve protocol generation, we propose the "Sketch-and-Fill" paradigm, which separates analysis, structuring, and expression to ensure each step is explicit and verifiable. Complementing this, the structured component-based reward mechanism evaluates step granularity, action order, and semantic fidelity, aligning model optimization with experimental reliability. Building on these components, we develop Thoth, trained through a staged Knowledge-to-Action process that progresses from knowledge acquisition to operational reasoning and ultimately to robust, executable protocol generation. Across multiple benchmarks, Thoth consistently surpasses both proprietary and open-source LLMs, achieving significant improvements in step alignment, logical sequencing, and semantic accuracy. Our approach paves the way for reliable scientific assistants that bridge knowledge with experimental execution. All data, code, and models will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。