用信息论方法量化大模型的智能水平,看它如何随上下文变化处理不确定性。
Measuring and Analyzing Intelligence via Contextual Uncertainty in Large Language Models using Information-Theoretic Metrics
- 基于熵衰减曲线,构建模型在不同上下文长度下的预测不确定性图谱。
- 发现模型规模和文本复杂度共同决定其熵衰减模式,且具有稳定性。
- 提出信息增益跨度指标,可一键评估模型内部信息处理能力优劣。
大语言模型在众多任务基准上表现优异,但其成功背后的机制仍不清晰。我们不再只关注模型能做什么,而是探究它们如何处理信息。本文提出一种无需任务依赖的方法,为任意模型构建定量认知轮廓。该轮廓基于熵衰减曲线——即模型在上下文长度增加时,归一化预测不确定性的变化趋势。在多个顶尖大模型与多样文本上,曲线展现出独特且稳定的特征,这些特征受模型规模和文本复杂度共同影响。我们进一步提出信息增益跨度(IGS)作为单一指标,综合反映熵衰减模式的优劣。这些工具为分析和比较现代人工智能系统的内部动态提供了理论基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel on many task-specific benchmarks, yet the mechanisms that drive this success remain poorly understood. We move from asking what these systems can do to asking how they process information. Our contribution is a task-agnostic method that builds a quantitative Cognitive Profile for any model. The profile is built around the Entropy Decay Curve -- a plot of a model's normalised predictive uncertainty as context length grows. Across several state-of-the-art LLMs and diverse texts, the curves expose distinctive, stable profiles that depend on both model scale and text complexity. We also propose the Information Gain Span (IGS) as a single index that summarises the desirability of a decay pattern. Together, these tools offer a principled way to analyse and compare the internal dynamics of modern AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。