arXiv:2602.10009cs.AIcs.HC2026-02被引 1

将仿真轨迹转化为高层结构模式,提升大模型对物理系统的理解能力

Discovering High Level Patterns from Simulation Traces

  • 通过程序合成自动识别仿真中的高层结构模式
  • 在物理基准测试中显著提升自然语言推理效果
  • 适合需要可解释性物理建模的科研与工程场景

大型语言模型(LLMs)难以可靠地推理特定物理系统。尽管赋予其物理知识已有良好进展,但可解释性与验证仍是挑战。一种新兴方法是工具化:让LLM调用物理模拟器,利用仿真轨迹作为上下文进行验证。但该方法因轨迹包含大量细粒度数值与语义数据而扩展性差。本文表明,将仿真轨迹转换为“高层”结构模式的稀疏表示,能更有效促进LLM的理解。我们提出一种无监督学习方案,通过程序合成实现此转换或标注。学习结果生成一套程序库,作为模式检测器,将仿真轨迹转为稀疏的带注释模式序列。检测模式可由人类专家通过字符串标签引导(如刚性碰撞、拉伸弹簧等)。在近期物理基准测试中,我们证明此类注释表示更利于自然语言推理特定物理系统。合成程序作为透明可解释函数,将系统状态映射至稀疏高效的注释空间。作为示例应用,我们展示了如何将自然语言目标转化为奖励程序,以最大化寻找解决方案。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are unable to reliably reason about specific physical systems. Attempts to imbue LLMs with knowledge of the necessary physics concepts have shown great promise, but explainability and validation remain open challenges. An emerging alternative is tooling, where LLMs can query physical simulators and use the resulting simulation traces as context for validation. This approach suffers from poor scalability since simulation traces contain large volumes of fine-grained numerical and semantic data. We show that translating simulation traces to a sparse representation of "high-level" structural patterns leads to more effective interpretation by LLMs. We propose an unsupervised learning scheme to perform this translation, or annotation, via program synthesis. Our learning results in a library of programs that act as pattern detectors which can translate simulation traces to sparse, annotated pattern sequences. The detected patterns may optionally be guided by human experts via string labels (rigid collision, stretching spring, etc.). We show, using a recent physics benchmark, that such annotated representations are more amenable to natural language reasoning about specific physical systems. The synthesized programs serve as transparent, explainable functions that map system states to a sparse and efficient annotation space. As an example application, we show how goals within physical systems that are specified in natural language may be converted to reward programs which are maximized to find solutions.

物理推理程序合成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。