用内部表示发现大模型与评测集的隐藏短板。
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
- 基于稀疏自编码器激活,按概念细粒度检测模型弱点。
- 自动发现模型曾被文献记录的缺陷及新短板。
- 适合评估者和评测集设计者改进基准体系。
大语言模型的评估高度依赖标准化基准。这些基准虽提供有用的整体指标,却可能掩盖模型在特定子领域的能力不足(模型差距)以及基准本身覆盖不均(基准差距)。为此,我们提出一种新方法,利用稀疏自编码器的概念激活,自动识别细粒度的模型与基准差距。该方法基于模型内部表征进行评估,便于跨基准比较。我们在五个主流开源模型和十余个基准上进行了验证。结果表明,该无监督自动方法成功复现了已有文献中记载的模型短板(如迎合性问题),并发现了新的模型短板;同时自动揭示了基准应涵盖但缺失的核心概念。该‘能力缺口’方法可作为现有基准的补充,提供模型行为的概念级分解,并帮助基准开发者优化设计。代码已公开:https://competency-gaps.github.io。
原文摘要 · Abstract (English)
The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can obscure (i) particular sub-areas where the models are weak ("model gaps") and (ii) imbalanced coverage in the benchmarks themselves ("benchmark gaps"). To automatically uncover both types of gaps, we propose a simple new method using concept activations from sparse autoencoders, to identify fine-grained gaps on a per-concept basis. The method also benefits from grounding evaluation in the model's internal representations, as well as easy comparison across benchmarks. We applied the method to five popular open-source models and more than a dozen benchmarks, as illustrative examples. As validation of the approach, we found that our automatic, unsupervised method was able to recover model gaps that have been previously documented in the literature (e.g. relating to sycophancy), in addition to identifying novel model gaps. We were also able to automatically uncover benchmark gaps: core concepts that should fall within the scope of a given benchmark. Our "competency gaps" method can be used to complement existing benchmarks, by providing a concept-level decomposition of model behavior, and by helping benchmark developers iterate upon benchmark design. Code is available at https://competency-gaps.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。