arXiv:2501.12011cs.CL2025-01综述被引 17

无需人工参考文本,自动评估生成文本质量的综合指南

Reference-free Evaluation Metrics for Text Generation: A Survey

  • 不依赖人工参考句,通过内在统计特征评估生成文本
  • 涵盖对话、摘要等多任务场景,适用于无标准答案的生成任务
  • 适合研究者快速了解主流无参考评估方法及其适用场景

自然语言生成系统已发展出多种自动评估指标。目前最常见的是基于参考文本的方法,即比较模型输出与人工撰写的标准参考句。然而,构建参考文本成本高,且在某些任务(如对话回复生成)中难以实现。近年来,多种无需参考的评估指标被提出。本综述全面覆盖各类自然语言生成任务,系统分析了常用方法的应用范围及在模型评估以外的其他用途。最后,指出了未来研究的若干有前景方向。

原文摘要 · Abstract (English)

A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's output with gold-standard references written by humans. However, it is expensive to create such references, and for some tasks, such as response generation in dialogue, creating references is not a simple matter. Therefore, various reference-free metrics have been developed in recent years. In this survey, which intends to cover the full breadth of all NLG tasks, we investigate the most commonly used approaches, their application, and their other uses beyond evaluating models. The survey concludes by highlighting some promising directions for future research.

文本生成评估方法无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。