arXiv:2608.25189cs.LGcond-mat.soft2026-08

让大模型像物理学家一样看数据,显著提升方程发现准确率。

What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery

论文配图:What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery
图 1 · 摘自论文原文
  • 将场数据转化为物理学家常用的可观测量输入模型
  • 在模拟数据上方程恢复准确率接近翻倍(提升至原3倍)
  • 无需训练、计算成本极低,适合与实验同步演化

理解分子相互作用如何决定宏观行为是分子科学的核心挑战。但传统理论构建难以跟上现代实验产生的海量数据。大语言模型为自动化理论构建提供了新路径,但时空场数据无法直接输入提示词。现有方法仅通过拟合得分来学习数据。本文提出数据解读阶段,将场数据转化为理论家会关注的物理量,并直接输入模型。在模拟场基准测试中,该方法使方程恢复准确率接近翻倍,计算开销可忽略,且无需任何训练。通过让大模型像物理学家一样读取场数据,数据解读为自动化场论构建提供了可行方案,可随实验同步演进。

原文摘要 · Abstract (English)

Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences. However, conventional theory building cannot keep pace with the vast datasets modern experimentation routinely produces. Large language models offer a promising route to automating theory construction, but a spatiotemporal field cannot be directly placed in a prompt. Existing models generally learn about the data only through a score measuring how well each proposal fits it. Here we introduce data interpretation, a stage that measures the field into the quantities a theorist would consult and supplies them to the model as a direct input. On a benchmark of simulated fields, interpretation nearly triples the accuracy of recovered equations relative to showing the raw data, at negligible computational cost and without any training. By allowing a language model to read field data as a theorist does, data interpretation offers a practical route to automated field theory construction that can coevolve with experimentation.

PDE发现大模型物理信息数据表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。