arXiv:2412.18011cs.CL2024-12被引 1

用结构化输出测试大模型推理能力,防作弊且成本低。

StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs

  • 通过规则引擎评估模型生成结构化输出的准确性
  • 17个主流模型在多领域任务中仍表现受限
  • 适合需要客观评测推理能力的研究者使用

大语言模型的快速进展亟需可靠、无偏且可扩展的评估方法。然而,人工标注成本高,基于模型的评估易受风格偏差影响,基于目标答案的基准又存在数据污染和作弊风险。为此,我们提出 StructTest,一种新型基准,通过检验模型对组合式指令的理解与结构化输出生成能力,提供一种无偏、低成本、难作弊的评估框架。评估采用基于规则的确定性评判器,可轻松扩展至新任务与数据集。我们在摘要、代码、HTML 和数学等多个领域测试了结构化输出,并评估了17个主流大模型,结果表明即使是最先进的 Deepseek-V3/R1 和 GPT-4o 仍面临挑战,证明 StructTest 可作为衡量推理能力的稳健代理。我们认为 StructTest 为实现客观、全面的模型评估提供了关键且互补的方法。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) demands robust, unbiased, and scalable evaluation methods. However, human annotations are costly to scale, model-based evaluations are susceptible to stylistic biases, and target-answer-based benchmarks are vulnerable to data contamination and cheating. To address these limitations, we propose StructTest, a novel benchmark that evaluates LLMs on their ability to follow compositional instructions and generate structured outputs, providing an unbiased, cost-effective, and difficult-to-cheat evaluation framework. Assessments are conducted deterministically using a rule-based evaluator, which can be easily extended to new tasks and datasets. By testing structured outputs across diverse domains including Summarization, Code, HTML, and Math, and evaluating 17 popular LLMs, we demonstrate that StructTest remains challenging even for top-performing models like Deepseek-V3/R1 and GPT-4o, establishing it as a robust proxy for measuring reasoning capabilities. We believe StructTest offers a critical and complementary approach to achieving objective and comprehensive model evaluation.

大模型评估结构化输出推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。