arXiv:2505.22830cs.CLcs.AI2025-05EMNLP被引 11

用大模型生成测评数据,虽省钱但难度下降,可能让评估失真。

What Has Been Lost with Synthetic Evaluation?

  • 用大模型生成阅读理解题,成本仅为人工标注的几分之一。
  • 生成题在逻辑有效性上达标,但对大模型来说明显更简单。
  • 适合关注评测数据质量、避免被模型“骗”的研究者参考。

大型语言模型(LLMs)正被广泛用于数据生成,但构建评估基准的门槛也随之提高。基准需针对特定现象、杜绝捷径利用并具备挑战性。本文通过两个案例研究,考察大模型能否通过生成推理类文本基准来满足这些要求,并与人工众包创建的高质量数据集进行对比。具体评估了两个高水准阅读理解数据集——CondaQA(测试否定推理)和DROP(测试数量推理)——的大模型生成版本。结果表明,提示大模型可低成本生成符合注释指南的有效题型,但其对大模型的挑战性远低于人工创建版本。这一发现揭示了用大模型生成评估数据时可能丢失的质量维度,呼吁重新审视该日益普遍的方法在基准构建中的直接应用。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for data generation. However, creating evaluation benchmarks raises the bar for this emerging paradigm. Benchmarks must target specific phenomena, penalize exploiting shortcuts, and be challenging. Through two case studies, we investigate whether LLMs can meet these demands by generating reasoning over-text benchmarks and comparing them to those created through careful crowdsourcing. Specifically, we evaluate both the validity and difficulty of LLM-generated versions of two high-quality reading comprehension datasets: CondaQA, which evaluates reasoning about negation, and DROP, which targets reasoning about quantities. We find that prompting LLMs can produce variants of these datasets that are often valid according to the annotation guidelines, at a fraction of the cost of the original crowdsourcing effort. However, we show that they are less challenging for LLMs than their human-authored counterparts. This finding sheds light on what may have been lost by generating evaluation data with LLMs, and calls for critically reassessing the immediate use of this increasingly prevalent approach to benchmark creation.

评测数据大模型生成评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。