arXiv:2505.24731cs.CL2025-05ACL被引 6

用电路稳定性评估大模型泛化能力,更可靠。

Circuit Stability Characterizes Language Model Generalization

  • 通过电路一致性衡量模型推理过程的稳定性。
  • 实验证明电路不稳时模型泛化能力下降。
  • 适合关注模型可解释性与泛化关系的研究者。

大规模语言模型的能力评估日益困难。先进模型快速迭代导致基准测试饱和,而设计更具挑战性的数据集又耗时费力。受机制可解释性研究启发,我们提出以电路稳定性作为评估模型性能的新方法。电路稳定性指模型在不同输入下保持一致推理路径(即电路)的能力。我们对电路稳定性与等价性进行了数学形式化,并通过三个案例研究,实证表明电路稳定性及其缺失能有效刻画并预测模型泛化行为的不同方面。本方法为严谨关联模型泛化性与可解释性迈出了关键一步。

原文摘要 · Abstract (English)

Extensively evaluating the capabilities of (large) language models is difficult. Rapid development of state-of-the-art models induce benchmark saturation, while creating more challenging datasets is labor-intensive. Inspired by the recent developments in mechanistic interpretability, we introduce circuit stability as a new way to assess model performance. Circuit stability refers to a model's ability to apply a consistent reasoning process-its circuit-across various inputs. We mathematically formalize circuit stability and circuit equivalence. Then, through three case studies, we empirically show that circuit stability and the lack thereof can characterize and predict different aspects of generalization. Our proposed methods offer a step towards rigorously relating the generality of models to their interpretability.

语言模型可解释性泛化能力电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。