arXiv:2607.20058cs.AIcond-mat.mes-hall2026-07

发现开源模型能通过状态变换隐含材料机制,且物理规律可被精准读取。

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

论文配图:Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
图 1 · 摘自论文原文
  • 用雅可比透镜和状态几何分析模型内部表示
  • 9/10机制家族可无标签识别,39/40方向定律正确排序
  • 适合研究AI可解释性与科学推理的学者

大型语言模型虽能回答科学问题,但其输出未必反映对物理规律的理解。本文在开源模型google/gemma-4-E4B-it中发现材料科学机制信息具有三种可分离表征形式:概念存在于单个隐藏状态中,本构关系由状态间的受控变换体现,特定内部表示则因果性地决定工程答案。研究结合直接与雅可比词典读出、无选项状态几何、60条定律的反事实基准及因果干预。在50个未见材料描述中,三个独立拟合的雅可比透镜重现了概念排名;两种读出方式生成的无目标词集成功盲识9/10机制族。另设72提示基准显示机制特异的状态邻域,但图审计表明此结构亦可由数值比较解释。进一步对比仅输入方向相反的同构提示,发现状态移动遵循给定本构律者达60个冻结关系中的39/40,而词汇控制接近随机。双向干预在12组匹配案例中使答案概率向物理合理结果偏移;反事实状态块在不同机制与答案格式间传递对立决策信号。因此,物理关系更显现在受控状态变化中而非绝对状态本身。

原文摘要 · Abstract (English)

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

科学推理可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。