单模型随机采样无法揭示知识盲区,只有多样模型集才能发现跨问题关联性。
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

- 用温度采样与多模型集成对比,检验语言模型的不确定性表达能力。
- 单模型在100次采样中仅1个显著维度,而24模型集成有4个显著维度。
- 适合关注模型可信度评估与不确定性建模的研究者阅读。
当语言模型在重复运行中给出不同答案时,这种变化是否反映了它不知道的内容?自一致性方法通过多数投票将变化转化为每个问题的不确定性估计。但同样的变化能否揭示跨问题的结构——即相关问题是否共同翻转,如同多样化集成模型那样?我们在相同问题上比较两种策略:单模型在τ=1下运行100次,与24个语言模型各运行一次(τ=0)组成的集成。采用Marchenko–Pastur随机矩阵检验区分信号与采样噪声。在任意单一模型中,五个模型族和三个基准(MMLU、HellaSwag、GSM8K)下,最多只有一个维度超出噪声水平;而在集成中,四个特征值显著高于噪声,而匹配难度的伯努利零模型在500次蒙特卡洛抽样中最多出现一个。自一致性提供了准确的单题不确定性,但无法检测到跨问题结构;只有多样化集成才能揭示模型未知之处。
原文摘要 · Abstract (English)
When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation into a per-question uncertainty estimate via majority voting. But does the same variation reveal cross-question structure -- related questions flipping together, the way a diverse ensemble does? We compare two regimes on the same questions: one model run $100$ times at $τ=1$ versus an ensemble of $24$ LLMs run once each at $τ=0$. A Marchenko--Pastur random-matrix test separates signal from sampling noise on both sides. Within any single model, at most one dimension rises above noise across five families and three benchmarks (MMLU, HellaSwag, GSM8K). Across the ensemble, four eigenvalues clear the noise edge, while a matched-difficulty Bernoulli null produces at most one in $500$ Monte Carlo draws. Self-consistency gives accurate per-question uncertainty but no detectable cross-question structure; only a diverse ensemble surfaces what a model does not know.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。