arXiv:2604.24076cs.AIcs.CL2026-04被引 1

用热力学类比分析大模型在不确定性下的稳定性表现

An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress

  • 构建融合任务效用与内部结构的综合稳定性评分
  • 相比基线平均提升0.0299,高熵条件下优势更明显
  • 适合关注AI可靠性与安全评估的研究者使用

随着大语言模型在高风险场景中广泛应用,仅依赖整体准确率的评估方法已难以刻画系统可靠性。本文提出一种受热力学启发的建模框架,用于分析大模型在不确定性和扰动下的输出稳定性。该框架引入综合稳定性得分,整合任务效用、熵(作为外部不确定性度量)以及两个内部结构代理变量:内部整合度与对齐反射能力。这些量不被视作物理变量,而是可解释的抽象,用以捕捉内部结构如何调节无序对模型行为的影响。基于IST-20基准协议及元数据,我们在四个主流大模型上分析了80个模型-场景观测结果。所提框架的稳定性得分均高于简化版效用-熵基线,平均提升0.0299(95%置信区间:0.0247–0.0351)。该增益在高熵条件下更为显著,表明框架能捕捉到不确定性非线性衰减的特征。本文不主张建立根本物理定律或完整机器伦理理论,而是贡献一个紧凑且可解释的统一评估视角,连接不确定性、性能与内部结构。该框架旨在补充现有评测方法,支持人工智能安全、可靠性与治理的持续讨论。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in high-stakes and operational settings, evaluation strategies based solely on aggregate accuracy are often insucient to characterize system reliability. This study proposes a thermodynamic inspired modeling framework for analyzing the stability of LLM outputs under conditions of uncertainty and perturbation. The framework introduces a composite stability score that integrates task utility, entropy as a measure of external uncertainty, and two internal structural proxies: internal integration and aligned reective capacity. Rather than interpreting these quantities as physical variables, the formulation is intended as an interpretable abstraction that captures how internal structure may modulate the impact of disorder on model behavior. Using the IST-20 benchmarking protocol and associated metadata, we analyze 80 modelscenario observations across four contemporary LLMs. The proposed formulation consistently yields higher stability scores than a reduced utilityentropy baseline, with a mean improvement of 0.0299 (95% CI: 0.02470.0351). The observed gain is more pronounced under higher entropy conditions, suggesting that the framework captures a form of nonlinear attenuation of uncertainty. We do not claim a fundamental physical law or a complete theory of machine ethics. Instead, the contribution of this work is a compact and interpretable modeling perspective that connects uncertainty, performance, and internal structure within a unied evaluation lens. The framework is intended to complement existing benchmarking approaches and to support ongoing discussions in AI safety, reliability, and governance.

大模型评估稳定性分析信息几何AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。