arXiv:2503.05551cs.LG2025-03被引 5

用加权评估法提升大模型性能区分度,解决基准测试饱和问题

Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance

  • 基于推理深度和难度动态分配权重,融合答案与思维链正确性
  • 在ARC-Challenge上将模型区分度从17%提升至46%
  • 适合需要精细评估大模型推理能力的研究者使用

现有基准测试因数据污染和大模型能力提升而趋于饱和,难以区分模型性能。本文提出EMDM(增强模型区分度度量),通过融合最终答案与思维链(CoT)正确性,并根据样本的复杂性和推理深度动态赋权,实现更细致的评估。在无引导(模型未接触测试样本)和有引导(已知目标答案)两种设置下,利用模型表现差异优化权重分配。相比仅依赖精确匹配(EM)的基准,在ARC-Challenge数据集上,EMDM将模型区分度从17%提升至46%,显著增强了对模型推理与知识需求的辨别能力。

原文摘要 · Abstract (English)

Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces EMDM (Enhanced Model Differentiation Metric), a novel weighted metric that revitalizes benchmarks by enhancing model separation. EMDM integrates final answer and Chain-of-Thought (CoT) reasoning correctness, assigning weights based on the complexity and reasoning depth required to solve a given sample in the evaluation data. Using a baseline LLM in two setups-Unguided, where the model has no prior exposure to test samples, and Guided, where the model has prior knowledge of the desired answer-EMDM distinguishes instances of varying difficulty. The CoT and answer correctness from these setups inform an optimization objective for weight assignment, resulting in a more nuanced evaluation of model performance. Compared to the exact match (EM) metric, which achieves 17% separation on ARC-Challenge, EMDM achieves 46%, demonstrating its effectiveness in differentiating models based on reasoning and knowledge requirements.

大模型评估推理能力加权度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。