构建可复现的模型能力评估体系,用多基准数据验证表示学习效果。
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

- 基于13,427篇论文构建182个能力簇,覆盖94项能力
- 46,149条经审计的探测文本支持跨基准验证
- 揭示读出方法与聚合策略影响评估结果,适合模型能力研究者
表示工程旨在读取和引导大语言模型的能力方向,但现有方法多依赖论文特有的合成数据,导致测量结果难以比较或复现,可能反映的是表面模式而非真实能力。我们提出RepBench,一个基于基准的、用于能力对齐表示探测的数据层。通过爬取13,427篇基准论文,构建了13类共182个能力簇;收集353个公开基准数据集,获得46,149条经审核的探测文本,覆盖94项能力,每项能力均得到至少两个独立基准支持。该多基准设计降低了对单一来源的依赖:单文本向量无自然聚类结构,而基准聚合后的能力向量在所有12个评估模型上均在少量聚类数时达到内部聚类最优,且与人类分类一致性低。跨基准迁移评估显示,在12个模型上,差异均值在10个模型上表现最佳,逻辑回归则在最多的能力-模型组合中胜出。这一分歧表明,读出方法与聚合标准是重要的评估维度。整个流程、语料库及评估代码已开源,形成可复用的闭环工作流。
原文摘要 · Abstract (English)
Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and may reflect surface patterns rather than capabilities. We present RepBench, a benchmark-grounded data layer for capability-aligned representation probing. Crawling 13,427 benchmark papers yields a taxonomy of 182 capability clusters in 13 families; harvesting 353 public benchmark datasets yields 46,149 audited probe texts covering 94 capabilities, each supported by at least two independent benchmarks. This multi-benchmark design reduces dependence on any single source: raw per-text vectors exhibit no natural cluster granularity, whereas benchmark-pooled capability vectors show an interior clustering optimum at a small number of clusters on all 12 evaluated models, with low agreement to the human taxonomy. Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells. This disagreement shows that the readout method and aggregation criterion are meaningful evaluation dimensions. The pipeline, corpus, and evaluation code are released as a reusable closed-loop workflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。