arXiv:2604.19341cs.LGcs.AI2026-04被引 10

通过结构化方法提升AI科学发现效率,实现多领域突破。

Structured Scaling of AI Discovery Across Diverse Scientific Domains

  • 设计新框架让AI在不同路径中高效试错与复用成果
  • 28个跨领域问题中多项性能超越此前纪录,最高降本24.5%
  • 适合关注AI辅助科研、算法优化与长期探索的研究者

科学发现常需反复提出、测试和优化方案。语言模型可参与这一过程,但盲目增加尝试次数无法保证进展:并行搜索易重复,迭代优化可能陷入低效路径。核心挑战不仅在于扩大规模,更在于结构化地组织规模,使评估信号随时间累积。本文提出SimpleTES(Simple Test-time Evaluation-driven Scaling)框架,聚焦于结构化扩展AI发现循环,统筹独立轨迹、迭代优化、局部候选选择及已评估历史的有选择性复用。基于科学共同体的结构特征,仅使用单一开源GPT-OSS模型,在量子物理、天文学、生物学、人工智能与数学等28个开放性问题上取得新最优解,包括量子电路编译开销降低24.5%、深空轨道推进成本降低最多23%、套索路径求解器速度提升2.17倍、全脑神经活动预测误差降低8.5%、最快报告的TriMul内核,以及超越人类和现有AI记录的新数学构造。进一步通过后训练赋予每轮尝试其所在轨迹的最终结果,显著提升数学问题上的训练与泛化表现。这些成果确立了结构化扩展作为推动AI科学发现的通用机制。

原文摘要 · Abstract (English)

Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions. Language models can increasingly participate in these loops, but simply generating more attempts does not ensure progress: parallel searches may duplicate one another and iterative refinement may become trapped in poor directions. The central challenge is therefore not only to scale AI-driven discovery, but to structure that scaling so that evaluation signals compound over time. Here we introduce SimpleTES (Simple Test-time Evaluation-driven Scaling), a framework that focuses on the structured scaling of AI discovery loops, organizing evaluator queries across independent trajectories, iterative refinement, local candidate selection, and the selective reuse of evaluated histories. Drawing on structural features of scientific communities, SimpleTES uses a single open-source GPT-OSS model to establish new state-of-the-art solutions across 28 open-ended problems in diverse scientific domains ranging from quantum physics and astronomy to biology, AI, and mathematics. These include a 24.5% reduction in quantum circuit compilation overhead, up to 23% lower propulsive cost for deep-space trajectories, a 2.17x faster lasso-path solver, an 8.5% lower-error whole-brain neural-activity predictor, the fastest reported TriMul kernel, and new mathematical constructions beyond prior human or AI records. We further post-train the model for long-horizon discovery by assigning each attempt the final outcome of the trajectory it helped produce. This improves performance on both training and held-out mathematics problems, further advancing the frontier. Together, these results establish structured scaling as a general mechanism for advancing AI scientific discovery.

科学发现语言模型结构化扩展跨领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。