arXiv:2508.15218cs.CL2025-08EMNLP被引 9

研究发现选择性使用检查清单能提升生成任务自动评估效果。

Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?

  • 对比六种生成方法,测试检查清单在不同模型规模下的表现
  • 选择性使用检查清单在成对比较中显著提升评估准确率
  • 即使相关性低的条目也反映人工标准,暴露人类评价不一致

使用大语言模型进行生成任务的自动评估面临标准模糊的问题。尽管自动检查清单生成具有潜力,但其实际效用尚未充分探索。本文研究检查清单应全量使用还是选择性使用,采用六种方法生成检查清单,评估其在八种不同模型规模下的有效性,并识别与人工评价相关性高的条目。在成对比较和直接打分任务中发现,选择性使用检查清单在成对设置中更优,而在直接打分中效果不一致。分析显示,即便某些条目与人工评分相关性较低,仍体现人工写作标准,暗示人工评价存在内在不一致性。研究强调需明确定义客观评估标准,以指导人工与自动评估。

原文摘要 · Abstract (English)

Automatic evaluation of generative tasks using large language models faces challenges due to ambiguous criteria. Although automatic checklist generation is a potentially promising approach, its usefulness remains underexplored. We investigate whether checklists should be used for all questions or selectively, generate them using six methods, evaluate their effectiveness across eight model sizes, and identify checklist items that correlate with human evaluations. Through experiments on pairwise comparison and direct scoring tasks, we find that selective checklist use tends to improve evaluation performance in pairwise settings, while its benefits are less consistent in direct scoring. Our analysis also shows that even checklist items with low correlation to human scores often reflect human-written criteria, indicating potential inconsistencies in human evaluation. These findings highlight the need to more clearly define objective evaluation criteria to guide both human and automatic evaluations. \footnote{Our code is available at~https://github.com/momo0817/checklist-effectiveness-study

自动评估生成模型检查清单

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。