用图结构增强推理,让AI生成可追溯的材料科学假设。
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

- 将推理分阶段建模,结合符号关系与语言生成,提升逻辑清晰度。
- 在100个开放问题上,推理可追溯性提升40%-65%,语义多样性提高2-3倍。
- 适合需要可解释性、可复现性的材料设计与科学发现研究者。
加速材料发现需要能通过多步、领域约束推理生成科学有效假设的AI系统。标准大模型常产生流畅但难以追溯的回答,难以判断最终结论是否由连贯中间推理支持。我们开发了基于图原生推理的Graph-PRefLexOR模型,采用分组相对策略优化(GRPO)微调,将推理划分为机制探索、图构建、模式提取和假设合成等显式阶段。该设计将神经语言生成与符号关系结构结合,实现因果关系的构建、检查与复用。在材料科学与力学文献中的100个开放问题上,Graph-PRefLexOR相比基线模型性能提升40%-65%,尤其在推理可追溯性方面改善显著。嵌入分析显示其语义探索范围更广,语义多样性约为基线的2-3倍。语义回溯与层间隐状态分析表明,结构化推理与最终答案对齐更强。测试时图扩展实验发现,额外计算主要促进有限语义空间内的长程概念重组,而非单纯扩大语义覆盖。结果确立图原生强化学习是实现材料设计等科学应用中可解释AI系统的可行路径。
原文摘要 · Abstract (English)
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. This design links neural language generation with symbolic relational structure, enabling causal connections to be constructed, inspected, and reused. On 100 open-ended questions from materials science and mechanics literature, Graph-PRefLexOR achieves 40-65% improvements over corresponding base models, with the largest gains in reasoning traceability. Embedding analyses show broader semantic exploration and approximately 2-3 times greater semantic diversity than baselines. Semantic backtracking and layer-wise hidden-state analyses further show stronger alignment between structured reasoning and final answers. Finally, test-time graph expansion reveals that additional compute primarily increases long-range conceptual recombination within a bounded semantic space, rather than simply expanding semantic coverage. These results establish graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。