arXiv:2602.18008cs.LGcs.AI2026-02

测试大模型在神经机制建模中的表现,提出新框架提升建模稳定性与质量。

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

  • 设计树状代理框架,分步探索复杂建模空间。
  • 在三个科学领域上显著提升搜索稳定性和解的质量。
  • 适合需要高可靠性科学建模的科研人员使用。

大型语言模型(LLMs)在从数据构建机制模型方面展现出潜力。然而,现有评估多集中于简化场景,难以反映真实科学建模的复杂性。实际建模常涉及神经网络与机制模型的联合构建,导致搜索空间显著复杂化。为此,我们引入神经集成机制建模(NIMM)基准,评估LLM生成的神经集成机制模型在三个科学领域的表现。实验表明,现有基于LLM的方法在该复杂空间中探索能力有限,搜索稳定性与解的质量均不足。为应对这一挑战,我们提出NIMMGen——一种基于树结构引导的代理框架,通过分支级搜索实现多样化探索,并通过原子模型精炼提升解的质量。大量实验证明,NIMMGen在NIMM基准上达到当前最优性能,显著改善搜索稳定性与解决方案质量。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on simplified settings and fail to capture the complexity of real-world scientific modeling. In practice, such modeling often involves neural-integrated formulations, where a mechanistic model component and a neural network component are jointly constructed, leading to a significantly more complex search space. Motivated by this gap, we introduce the Neural-Integrated Mechanistic Modeling (NIMM) benchmark, which evaluates LLM-generated neural-integrated mechanistic models across three scientific domains. Experiments on NIMM reveal that existing LLM-based approaches struggle to effectively explore this complex space, resulting in limited search stability and solution quality. To address this challenge, we propose NIMMGen, a tree-guided agentic framework that enables diversified exploration via branch-level search and improves solutions through atomic model refinement. Extensive experiments demonstrate that NIMMGen achieves state-of-the-art performance on NIMM, significantly improving search stability and solution quality.

机制建模大模型代理框架科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。