arXiv:2512.23292cs.AIcs.LG2025-12中稿 · publication in npj…被引 1

用小型语言模型实现核反应堆控制的可靠决策,突破通用AI在物理系统中的局限。

Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control

  • 用物理仿真验证驱动策略优化,而非依赖感知推理
  • 3600万参数模型在10万样本下将误差波动降低500倍
  • 适合需要高可靠性的工业控制系统,不适用于故障或异常场景

当前通用基础模型在物理系统中的应用面临控制接口瓶颈:前沿视觉语言模型在基础物理任务中准确率仅50%-53%,行为类似近似猜测,虽语义合理却违背物理约束。安全关键控制需对执行动作的结果空间提供保障,而非参数空间的模仿。本文提出一种面向特定领域的基础模型路径:采用小型语言模型作为代理型物理智能(Agentic Physical AI),通过物理仿真验证驱动策略优化。我们在合成核反应堆场景上训练了一个360M参数模型,样本量从10³扩展至10⁵。规模扩大带来显著的、与工况相关的可靠性提升,在正常模拟条件下,方差降低约500倍,并消除超过10%的终端功率超限。尽管四类操作策略均等暴露,模型95%运行时间仅使用单一策略,且未使用强化学习或奖励工程。该表示可跨仿真器迁移,无需改变架构。我们将其定位为验证、监控和纵深防御体系中的决策组件,而非独立安全方案:当前表现仅证明单步任务在仿真中的闭环可靠性,尚未涵盖非正常运行、传感器故障或不确定性量化。

原文摘要 · Abstract (English)

The prevailing paradigm in AI for physical systems: scaling general-purpose foundation models toward universal multimodal reasoning, confronts a barrier at the control interface. Frontier vision-language models achieve only 50-53% accuracy on basic quantitative physics tasks, behaving as approximate guessers that preserve semantic plausibility while violating physical constraints. Safety-critical control demands outcome-space guarantees over executed actions, not parameter-space imitation. Here we present a pathway toward domain-specific foundation models through compact language models operating as Agentic Physical AI: policy optimization driven by physics-based simulator validation rather than perceptual inference. We train a 360M-parameter model on synthetic nuclear reactor scenarios scaled from 10^3 to 10^5 examples. Scaling produces strong, regime-dependent reliability gains under nominal simulated conditions, with variance collapse of approximately 500x and elimination of >10% terminal-power excursions on the sampled distribution. Despite balanced exposure to four actuation families, the model concentrates 95% of runtime execution on a single-bank strategy, without reinforcement learning or reward engineering. Representations transfer across simulators without architectural change. We position the system as a candidate decision component within a verification, monitoring, and defense-in-depth architecture, not as a stand-alone safety solution: the demonstrated behavior speaks to closed-loop reliability on a single-step task in simulation and does not yet address off-nominal operation, sensor faults, or uncertainty quantification.

物理AI核反应堆代理模型仿真验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。