提出新指标衡量大模型对输入扰动的内在稳定性,揭示传统方法忽略的预测脆弱性。
Beyond Confidence: The Rhythms of Reasoning in Generative Models
- 引入Token Constraint Bound(δ_TCB)量化模型内部状态抗扰能力
- δ_TCB与提示工程效果相关,在上下文学习中发现困惑度未捕捉的不稳定性
- 适用于评估和提升大模型生成结果的上下文鲁棒性
大型语言模型虽能力出众,但对输入上下文微小变化敏感,影响可靠性。传统指标如准确率和困惑度无法评估局部预测的鲁棒性,因归一化输出概率会掩盖模型内部状态对扰动的响应。本文提出新指标Token Constraint Bound(δ_{TCB}),量化模型在主导下一个词预测发生显著变化前可承受的最大内部状态扰动。δ_{TCB}与输出嵌入空间几何结构密切相关,揭示了模型内部预测承诺的稳定性。实验表明,δ_{TCB}与有效提示工程相关,并在上下文学习和文本生成中发现了困惑度未能察觉的关键预测不稳定性。该指标为分析和可能改进大模型预测的上下文稳定性提供了原则性、互补性方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit impressive capabilities yet suffer from sensitivity to slight input context variations, hampering reliability. Conventional metrics like accuracy and perplexity fail to assess local prediction robustness, as normalized output probabilities can obscure the underlying resilience of an LLM's internal state to perturbations. We introduce the Token Constraint Bound ($δ_{\mathrm{TCB}}$), a novel metric that quantifies the maximum internal state perturbation an LLM can withstand before its dominant next-token prediction significantly changes. Intrinsically linked to output embedding space geometry, $δ_{\mathrm{TCB}}$ provides insights into the stability of the model's internal predictive commitment. Our experiments show $δ_{\mathrm{TCB}}$ correlates with effective prompt engineering and uncovers critical prediction instabilities missed by perplexity during in-context learning and text generation. $δ_{\mathrm{TCB}}$ offers a principled, complementary approach to analyze and potentially improve the contextual stability of LLM predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。