让AI发现的物理模型通过真实物理规则检验,避免只看误差假象。
Physics-Audited Agentic Discovery in Scientific Machine Learning

- 先定验证标准,再搜索模型,每候选都检查物理合规性。
- 在瞬态弹性中,传统方法误响应未来载荷,新方法通过因果检验。
- 适合需要高可信物理建模的研究者,尤其力学与仿真领域。
在智能科学机器学习(SciML)中,大语言模型(LLM)代理可通过误差指标自动筛选代理模型。但低误差未必代表预测场满足关键物理规律,如边界条件、叠加原理、刚度缩放或因果性。本文提出物理审计式智能科学机器学习(PA-SciML),采用验证优先的工作流:固定评分评估器后,推导可审查、机器可检的物理要求,对每个训练候选模型输出进行逐项检查,并在指定输入范围或实测载荷历史区间内搜索高违反案例,无需参考解。仅在所有检查通过时才报告为已验证模型。启用后,工作流还加入训练前数值探针,并单次测试一个建模改动,记录其对得分的影响以供复用。在计算固体力学的数值示例中,静态弹性实验中,新方法选出的模型虽验证误差低于基线,但两者均通过常见线弹性检验;在瞬态弹道动力学实验中,误差基线模型平均误差相近,却因响应未来载荷历史而违反更严格的因果性检验,而新选模型通过全部规定检查。核心差异在于对每个候选模型输出的物理证据,而非更复杂的综合评分。
原文摘要 · Abstract (English)
In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, does not establish that the predicted fields satisfy the physics that matter for mechanics, such as boundary conditions, superposition, stiffness scaling, or causality. We introduce Physics-Audited Agentic SciML (PA-SciML), a verification-first workflow for agentic SciML discovery. The workflow fixes a scoring evaluator before search, derives reviewable machine-checkable physics requirements, checks each trained candidate on its outputs, and separately searches prescribed input ranges or measured load-history spans for high-violation cases without reference solution fields. A surrogate is reported as verified only under the stated checks. When enabled, the workflow also adds advisory numerical probes before training and tests one modeling change at a time to record which isolated edits are associated with score gains before reuse. In the reported computational-solid-mechanics numerical examples, the static elasticity run selects a surrogate with lower validation error than the error-only baseline while both selected models pass the common linear-elastic checks. In the transient elastodynamics run, an error-only baseline with similar mean error fails a stricter causality check by responding to future parts of the loading history, while the selected surrogate passes the stated checks. The main distinction is per-candidate physics evidence on predicted fields, not a richer aggregate score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。