从样本级剖析评测集,让大模型评估更精准
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

- 按五个维度逐样本分析评测数据,揭示内部差异
- 发现平均分掩盖了样本间巨大差异,不同任务需求迥异
- 可按需组合评测子集,适合关注推理或伦理的开发者
评测数据集是评估大语言模型的核心,但通常被视为整体任务,掩盖了单个样本间的显著差异。我们提出一种以数据集为中心的元评估框架,在认知与知识需求、语言与内容质量、任务属性、上下文以及伦理、安全与公平五个潜在维度上对评测数据进行样本级审计。应用该框架,我们对五个有影响力的基准——MMLU、ARC、WinoGrande、HellaSwag 和 TruthfulQA——进行了标注,揭示出显著的内部异质性,这些差异无法通过整体准确率反映。我们展示了如何利用这些标注实现基于标准的复合评测子集编排,支持对模型能力如推理深度或伦理敏感性的针对性评估。该方法将评测评估重新定义为数据集的自我审视,提供了一种系统化分析和重组现有评测集的方法,以更好满足多样化的评估需求。
原文摘要 · Abstract (English)
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. Cognitive and Knowledge Demands, 2. Language and Content Quality, 3. Task Properties, 4. Context, and 5. Ethics, Safety, and Fairness. Applying this framework, we annotate five influential benchmarks -- MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA -- revealing pronounced internal heterogeneity that is not captured by aggregate accuracy scores. We show how these annotations enable criterion-driven orchestration of composite benchmark subsets across datasets, supporting targeted evaluation of model capabilities such as Reasoning Depth or Ethical Sensitivity. This approach reframes benchmark evaluation as dataset introspection, providing a principled methodology for analyzing and re-composing existing benchmarks to better reflect diverse evaluation needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。