arXiv:2511.20798cs.LGcs.AI2025-11被引 6

用物理模型内部概念向量,可精准操控模拟中的物理行为。

Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model

  • 从物理仿真数据中提取激活向量,计算不同物理状态的差值作为概念方向。
  • 注入这些方向后,能有效诱导或消除模拟中的特定物理特征。
  • 证明科学基础模型具备对物理原理的泛化表征,适合科研与可控生成场景。

近期机制可解释性研究发现,大语言模型不仅包含具体实体的表征,还存在可被直接操纵的人类可理解的抽象概念和行为。然而,这一现象是否仅限于结构化数据(如语言、图像)训练的模型,还是基础模型的普遍特性仍不清楚。本文研究了一个面向物理领域的大型基础模型的内部表征。受启发于此前在大语言模型中识别复杂行为单一方向的工作,我们在模型对不同物理状态仿真数据的前向传播过程中提取激活向量,并计算两状态间的“差值”表示。这些差值张量构成了激活空间中的概念方向,编码了特定物理特征。通过在推理阶段将这些方向注入模型,我们实现了对预测结果的因果控制,例如在模拟中引入或移除特定物理特征。结果表明,科学基础模型学习到了物理原理的泛化表征,不依赖于仿真中的表面相关性。该研究为理解与控制科学基础模型开辟新路径,对人工智能驱动的科学发现具有重要意义。

原文摘要 · Abstract (English)

Recent advances in mechanistic interpretability have revealed that large language models (LLMs) develop internal representations corresponding not only to concrete entities but also distinct, human-understandable abstract concepts and behaviour. Moreover, these hidden features can be directly manipulated to steer model behaviour. However, it remains an open question whether this phenomenon is unique to models trained on inherently structured data (ie. language, images) or if it is a general property of foundation models. In this work, we investigate the internal representations of a large physics-focused foundation model. Inspired by recent work identifying single directions in activation space for complex behaviours in LLMs, we extract activation vectors from the model during forward passes over simulation datasets for different physical regimes. We then compute "delta" representations between the two regimes. These delta tensors act as concept directions in activation space, encoding specific physical features. By injecting these concept directions back into the model during inference, we can steer its predictions, demonstrating causal control over physical behaviours, such as inducing or removing some particular physical feature from a simulation. These results suggest that scientific foundation models learn generalised representations of physical principles. They do not merely rely on superficial correlations and patterns in the simulations. Our findings open new avenues for understanding and controlling scientific foundation models and has implications for AI-enabled scientific discovery.

物理建模概念控制基础模型因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。