用隐藏状态构建小样本核心集,高效估算大模型性能。
Learning More from Less: Unlocking Internal Representations for Benchmark Compression
- 通过对齐隐状态构造统一表征空间,生成代表性核心样本。
- 仅需10个源模型即可实现高精度性能预测,优于传统方法。
- 适合新发布基准的快速评估,尤其适用于数据稀缺场景。
大型语言模型(LLM)评估成本高昂,亟需高效替代方案。现有方法依赖从多个源模型的响应中估计项目特征,但在源模型数量少时统计不稳,尤其限制新发布基准的应用。本文认为离散正确性标签无法捕捉隐藏状态中的决策信息。提出RepCore方法,将异构隐藏状态对齐至统一潜在空间,构建代表性核心集。实验在五个基准、200余模型上验证,仅用10个源模型即实现更优的排名相关性和估计准确率。谱分析显示,对齐表示包含反映普遍响应倾向与任务特定推理模式的可分离成分。
原文摘要 · Abstract (English)
The prohibitive cost of evaluating Large Language Models (LLMs) necessitates efficient alternatives to full-scale benchmarking. Prevalent approaches address this by identifying a small coreset of items to approximate full-benchmark performance. However, existing methods must estimate a reliable item profile from response patterns across many source models, which becomes statistically unstable when the source pool is small. This dependency is particularly limiting for newly released benchmarks with minimal historical evaluation data. We argue that discrete correctness labels are a lossy view of the model's decision process and fail to capture information encoded in hidden states. To address this, we introduce RepCore, which aligns heterogeneous hidden states into a unified latent space to construct representative coresets. Using these subsets for performance extrapolation, RepCore achieves precise estimation accuracy with as few as ten source models. Experiments on five benchmarks and over 200 models show consistent gains over output-based baselines in ranking correlation and estimation accuracy. Spectral analysis further indicates that the aligned representations contain separable components reflecting broad response tendencies and task-specific reasoning patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。