发现大模型回答越不确定,探测模型表现越差。
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
- 用探测方法分析模型响应不确定性与性能的关系。
- 响应方差越大,重要特征越多,探测效果越差。
- 揭示了模型可解释性背后的内在机制,适合研究者参考。
探测技术在揭示大语言模型如何编码人类可理解概念方面展现出潜力,尤其在使用精心构建的数据集时。然而,决定数据集是否适合有效探测训练的因素尚不明确。本研究假设,探测性能反映了大模型生成响应及内部特征空间的特性。通过在一系列任务中对探测性能和模型响应不确定性进行定量分析,我们发现两者存在强相关性:探测性能提升始终对应响应不确定性降低,反之亦然。进一步从特征重要性角度深入分析,发现模型响应方差越高,重要特征集合越大,给探测模型带来更大挑战,常导致性能下降。此外,利用响应不确定性分析的洞察,我们识别出多个领域中模型表示与人类知识一致的具体案例,为大模型具备可解释推理提供了额外证据。
原文摘要 · Abstract (English)
Probing techniques have shown promise in revealing how LLMs encode human-interpretable concepts, particularly when applied to curated datasets. However, the factors governing a dataset's suitability for effective probe training are not well-understood. This study hypothesizes that probe performance on such datasets reflects characteristics of both the LLM's generated responses and its internal feature space. Through quantitative analysis of probe performance and LLM response uncertainty across a series of tasks, we find a strong correlation: improved probe performance consistently corresponds to a reduction in response uncertainty, and vice versa. Subsequently, we delve deeper into this correlation through the lens of feature importance analysis. Our findings indicate that high LLM response variance is associated with a larger set of important features, which poses a greater challenge for probe models and often results in diminished performance. Moreover, leveraging the insights from response uncertainty analysis, we are able to identify concrete examples where LLM representations align with human knowledge across diverse domains, offering additional evidence of interpretable reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。