用可验证代码评估大模型是否能精准执行复杂指令中的字符串匹配任务。
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
- 通过严格步骤指令和可验证代码,实现客观评测。
- 测试模型在指令遵循、数值计算和中间结果一致性上的表现。
- 适合评估前沿大模型的精确执行能力,尤其关注细节处理。
当前顶尖大语言模型已在许多基准上达到饱和,难以进一步区分性能。这凸显了需要更具挑战性的、可客观验证的评测基准。本文提出MCBench,一个用于评估大模型能否根据逐步指令执行字符串匹配自然语言处理指标的基准。与依赖主观判断或通用推理的先前基准不同,MCBench提供客观、确定且可代码验证的评估方式。该设计使我们能够系统性测试模型在指令遵循、数值计算及中间结果长程一致性方面的准确执行能力。为确保评估客观性,我们提供并行参考代码以验证模型输出的准确性。本研究包含三个评估指标和三种基准变体,旨在衡量大模型对复杂指令的理解细节。分析表明,MCBench是评估前沿大模型能力的有效且客观工具。
原文摘要 · Abstract (English)
Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation. This progress highlights the need for challenging benchmarks that provide objective verification. In this paper, we introduce MCBench, a benchmark designed to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions. Unlike prior benchmarks that depend on subjective judgments or general reasoning, MCBench offers an objective, deterministic and codeverifiable evaluation. This setup allows us to systematically test whether LLMs can maintain accurate step-by-step execution, including instruction adherence, numerical computation, and long-range consistency in handling intermediate results. To ensure objective evaluation of these abilities, we provide a parallel reference code that can evaluate the accuracy of LLM output. We provide three evaluative metrics and three benchmark variants designed to measure the detailed instruction understanding capability of LLMs. Our analyses show that MCBench serves as an effective and objective tool for evaluating the capabilities of cutting-edge LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。