arXiv:2604.11945cs.LGcs.AI2026-04

让不懂机器学习的人用自然语言自动生成地下流体模拟的高效代理模型。

AutoSurrogate: An LLM-Driven Multi-Agent Framework for Autonomous Construction of Deep Learning Surrogate Models in Subsurface Flow

论文配图:AutoSurrogate: An LLM-Driven Multi-Agent Framework for Autonomous Construction of Deep Learning Surrogate Models in Subsurface Flow
图 1 · 摘自论文原文
  • 基于大模型的多智能体协作,自动完成从数据到模型的全流程构建。
  • 仅需一句自然语言指令,即可生成满足精度要求的可部署模型。
  • 自动处理训练失败、精度不足等问题,适合无ML背景的领域科学家使用。

高保真地下流体数值模拟计算成本高昂,尤其在不确定性量化与数据同化等需多次查询的任务中。深度学习代理模型可显著加速正向模拟,但其构建需要大量机器学习专业知识——从网络架构设计到超参数调优,大多数领域科学家缺乏此类能力。此外,该过程高度依赖人工和经验性选择,成为制约深度学习代理广泛应用的关键障碍。为此,本文提出 AutoSurrogate,一个由大语言模型驱动的多智能体框架,使无机器学习背景的使用者可通过自然语言指令构建高质量的地下流体代理模型。给定仿真数据及可选偏好,四个专用智能体协同完成数据分析、模型库中的架构选择、贝叶斯超参数优化、模型训练及质量评估,并在出现数值不稳定或预测精度不达标时自主重启训练或切换架构。在三维地质碳封存建模任务中,以31个时间步的渗透率场预测压力与CO₂饱和度场,仅需一句自然语言指令,无需手动调整,AutoSurrogate即能超越专家设计基线与通用AutoML方法,展现出强大的实际应用潜力。

原文摘要 · Abstract (English)

High-fidelity numerical simulation of subsurface flow is computationally intensive, especially for many-query tasks such as uncertainty quantification and data assimilation. Deep learning (DL) surrogates can significantly accelerate forward simulations, yet constructing them requires substantial machine learning (ML) expertise - from architecture design to hyperparameter tuning - that most domain scientists do not possess. Furthermore, the process is predominantly manual and relies heavily on heuristic choices. This expertise gap remains a key barrier to the broader adoption of DL surrogate techniques. For this reason, we present AutoSurrogate, a large-language-model-driven multi-agent framework that enables practitioners without ML expertise to build high-quality surrogates for subsurface flow problems through natural-language instructions. Given simulation data and optional preferences, four specialized agents collaboratively execute data profiling, architecture selection from a model zoo, Bayesian hyperparameter optimization, model training, and quality assessment against user-specified thresholds. The system also handles common failure modes autonomously, including restarting training with adjusted configurations when numerical instabilities occur and switching to alternative architectures when predictive accuracy falls short of targets. In our setting, a single natural-language sentence can be sufficient to produce a deployment-ready surrogate model, with minimum human intervention required at any intermediate stage. We demonstrate the utility of AutoSurrogate on a 3D geological carbon storage modeling task, mapping permeability fields to pressure and CO$_2$ saturation fields over 31 timesteps. Without any manual tuning, AutoSurrogate is able to outperform expert-designed baselines and domain-agnostic AutoML methods, demonstrating strong potential for practical deployment.

代理模型大模型地下流体自动化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。