arXiv:2606.11657cs.LGcs.AI2026-06被引 2

用物理指标筛选特征,发现科学大模型内部机制与真实物理不完全对应。

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

论文配图:Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics
图 1 · 摘自论文原文
  • 通过稀疏自编码器探测模型层,用涡度筛选超2万特征
  • 不同剪切流设置下特征重复出现但结构断续,未对齐标准物理分解
  • 模型输出偏差与特定特征使用变化相关,揭示解释性挑战

生成式人工智能模拟器在已有强理论、基准和物理直觉的科学领域日益普及。这引出核心问题:当基础模型能复现已知连续介质动力学时,其内部机制是否符合物理规律,行为是否一致,以及如何关联模型的成功与失败?本文以Polymathic的Walrus模型为研究对象,采用基于物理原理的机械可解释性方法,对选定层应用稀疏自编码器(SAE),并利用涡度作为物理基础指标,应对超20,000个特征的筛选难题。以剪切流为简单测试场景,比较不同参数设置下的特征激活模式。结果显示特征呈现分段一致性,部分特征在相似角色中反复出现,但整体结构间歇且未能清晰映射至标准物理分解。同时,直接对比数值模拟与模型输出发现系统性差异,如能量或结构过度弥散或过度集中。我们进一步将部分差异归因于特定SAE特征的使用变化。本工作揭示了科学类基础模型的关键开放问题:如何稳健地优先选择机制上有意义的特征;如何区分稳定结构与分析伪影(包括单层及SAE局限);以及如何利用既有基准判断‘内部表征不同’是真正有信息量还是仅有效而已。

原文摘要 · Abstract (English)

Generative AI emulators are increasingly used in scientific domains where we already have strong theory, benchmarks, and physical intuition. This raises a central evaluation and interpretability question: when a foundation-style model can reproduce known continuum dynamics, what internal mechanism supports that behavior, is the internal behaviour consistent with known physics, and how does it relate to where the emulator succeeds or fails? We investigate a cross-domain foundation model for continuum dynamics, Walrus by Polymathic, using mechanistic interpretability guided by physical principles. We apply a sparse autoencoder (SAE) to probe a selected layer, and address the practical challenge of triaging a large feature set (over 20,000) using enstrophy as a physically grounded metric. As a deliberately simple testbed, we focus on shear flow and compare feature recruitment across multiple shear-flow setups, i.e. parameter values in the numerical simulation. Across setups we find evidence of piecewise consistency, with subsets of features recurring in similar roles, but this structure is intermittent and does not map cleanly onto standard physical decompositions. In parallel, direct comparisons between numerical simulation and the emulator reveal systematic output-level discrepancies, including regimes where energy/structures become too diffuse or too localized. We connect parts of these discrepancies to changes in specific SAE feature usage. Our work highlights open questions for scientific foundation models: how to robustly prioritize mechanistically meaningful features, how to separate stable structure from analysis artifacts (including single-layer and SAE limitations), and how to use established benchmarks to decide when "different" internal representations are genuinely informative rather than merely effective.

可解释性科学建模基础模型物理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。