AI能否发现物理问题的统计力学映射?实验揭示其潜力与局限。
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
- 用大模型构建试错修正的智能体,尝试从配分函数中发现可解结构。
- 数值验证能修复代码但无法确保正确识别物理模型或复杂度判断。
- 适合对物理建模、符号推理和AI辅助科研感兴趣的学者参考。
理论物理中的重要能力是识别新问题是否可转化为已知模型。本文将此能力作为AI智能体的任务:基于大语言模型的智能体能否从原始配分函数中发现可解的统计力学映射?为此,我们引入StatMechBench-v0基准,包含六个伊辛型问题,涵盖转移矩阵法、可消除规范无序及平面/Pfaffian结构。我们在多个大模型和问题表述下评估了一个简单的提出-验证-修正智能体。结果表明,数值反馈常能帮助智能体修复代码并恢复正确的配分函数,但智能体也可能在通过数值检验的同时错误识别底层可解类别或低估计算复杂性。这既揭示了当前大模型推理的局限性,也呼吁建立超越数值一致性的验证体系,例如符号检查与结构不变量。本研究为面向理论物理结构性发现的AI智能体提供了早期评估与设计方向。
原文摘要 · Abstract (English)
An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representation? To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and planar/Pfaffian structure. We evaluate a simple propose-verify-revise agent across multiple LLMs and problem phrasings. The results show that numerical feedback often helps agents repair code and recover correct partition functions. However, agents can also pass the numerical checks while misidentifying the underlying tractable class or understating computational complexity. This both reveals limitations in current LLM reasoning and calls for a verification stack that goes beyond numerical agreement, incorporating, for example, symbolic checks and structural invariants. Our study provides an early evaluation and design directions for AI agents aimed at structural discovery in theoretical physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。