arXiv:2410.08437cs.AIcs.CL2024-10ICLR被引 5

自动生成任务与答案,实现无需人工标注的LLM真值维护与推理能力评估。

Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks

  • 自动生成不同难度的任务和真值答案,支持持续评估更复杂的LLM。
  • 在翻译与逻辑推理任务上表现与多个基准高度一致,验证了有效性。
  • 适合缺乏人力标注资源或需动态更新评估的场景,如持续模型迭代。

本文提出AutoEval,一个面向形式化任务(如翻译中的真值维护、逻辑推理)的新型基准,用于规模化评估大语言模型。它是首个无需人工标注即可实现客观评估的范式:(a) 能通过自动生成不同难度的任务,评估日益复杂的模型;(b) 自动生成真值,避免依赖昂贵的人工标注;(c) 使用随机生成的数据集,防止后续模型对静态数据过拟合。实证分析表明,LLM在AutoEval上的表现与其在多种翻译与推理类基准上的性能高度相关,证明其在难以获取或更新人工标注数据的场景下具有重要价值。

原文摘要 · Abstract (English)

This paper presents AutoEval, a novel benchmark for scaling Large Language Model (LLM) assessment in formal tasks with clear notions of correctness, such as truth maintenance in translation and logical reasoning. AutoEval is the first benchmarking paradigm that offers several key advantages necessary for scaling objective evaluation of LLMs without human labeling: (a) ability to evaluate LLMs of increasing sophistication by auto-generating tasks at different levels of difficulty; (b) auto-generation of ground truth that eliminates dependence on expensive and time-consuming human annotation; (c) the use of automatically generated, randomized datasets that mitigate the ability of successive LLMs to overfit to static datasets used in many contemporary benchmarks. Empirical analysis shows that an LLM's performance on AutoEval is highly indicative of its performance on a diverse array of other benchmarks focusing on translation and reasoning tasks, making it a valuable autonomous evaluation paradigm in settings where hand-curated datasets can be hard to obtain and/or update.

大模型评估自动化测试推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。