arXiv:2606.04505cs.AI2026-06

让大模型理解科学模拟器的内部机制,提升决策透明度与可靠性。

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

论文配图:Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making
图 1 · 摘自论文原文
  • 构建共享结构化框架,显式表示模拟器的假设、变量与机制依赖。
  • 大模型基于机制生成有证据支持的解释,提升对模拟结果的理解。
  • 适用于高风险决策场景,如气候预测、医疗模拟等需可解释性的领域。

科学模拟器正越来越多地被集成到基于大模型的系统中,用于高风险的模拟驱动决策。然而,现有框架主要将大模型用于生成、校准或执行模拟器,将其视为黑箱接口,而非可推理的机制化系统。这导致当前方法无法识别、表征和推理模拟器行为背后的假设与机制,限制了透明性、可审计性和决策解释能力。本文提出 MechSim,一个以机制为基础的神经符号推理框架,用于可执行的科学模拟器。与以往主要处理静态符号结构的神经符号方法不同,MechSim 使大模型代理能够推理科学模拟器的机制、假设及执行行为。该框架通过共享的结构化模式捕获假设、变量、机制依赖关系和执行轨迹。在此基础上,大模型代理作为受约束的推理引擎,生成结构化且基于证据的解释,将模拟结果与其底层机制联系起来。我们在多个高风险领域评估该方法,结果表明其显著提升了机制级解释质量、模拟器分析能力以及下游决策的可靠性。

原文摘要 · Abstract (English)

Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making. However, existing frameworks primarily use LLMs to generate, calibrate, or execute simulators, treating them as black-box interfaces rather than as structured mechanistic systems that can be reasoned about. As a result, current approaches lack the ability to identify, represent, and reason about the assumptions and mechanisms underlying simulator behavior, limiting transparency, auditability, and decision justification. We introduce MechSim, a mechanism-grounded neuro-symbolic reasoning framework for executable scientific simulators. Unlike prior neuro-symbolic approaches that primarily reason over static symbolic structures, MechSim enables LLM agents to reason about the mechanisms, assumptions, and execution behavior of scientific simulators. Our framework represents simulators through a shared structured schema capturing assumptions, variables, mechanism dependencies, and execution traces. On top of this representation, LLM agents operate as constrained reasoning engines that generate structured, evidence-grounded explanations linking simulator outcomes to their underlying mechanisms. We evaluate our approach across multiple high-stakes domains and show that it improves mechanism-level explanation quality, simulator analysis, and downstream decision-making reliability.

科学推理大模型可解释性模拟系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。