用统一标准重评文本生成系统,发现多数性能被高估。
A Comparative Study of Controlled Text Generation Systems Using Level-Playing-Field Evaluation Principles

- 所有系统用相同数据和评估方法处理输出
- 多数系统重评后性能明显下降
- 适合关注评估公平性的研究者
近年来提出了多种受控文本生成(CTG)方法,但因使用不同数据集和评估方式,难以判断哪种方法最优。本文提出基于均等竞技场(LPF)的评估方法,统一标准化处理所有系统输出,并采用当前主流的评估方法与数据集进行公平比较。在该标准下重新评估一组代表性CTG系统,结果与原始报告差异显著,多数系统表现更差。这凸显了建立标准化、可复现评估流程的重要性。研究显示,缺乏统一标准可能导致性能宣称严重偏离真实能力。
原文摘要 · Abstract (English)
Background: Many different approaches to controlled text generation (CTG) have been proposed over recent years, but it is difficult to get a clear picture of which approach performs best, because different datasets and evaluation methods are used in each case to assess the control achieved. Objectives: Our aim in the work reported in this paper is to develop an approach to evaluation that enables us to comparatively evaluate different CTG systems in a manner that is both informative and fair to the individual systems. Methods: We use a level-playing-field (LPF) approach to comparative evaluation where we (i) generate and process all system outputs in a standardised way, and (ii) apply a shared set of evaluation methods and datasets, selected based on those currently in use, in order to ensure fair evaluation. Results: When re-evaluated in this way, performance results for a representative set of current CTG systems differ substantially from originally reported results, in most cases for the worse. This highlights the importance of a shared standardised way of assessing controlled generation. Conclusions: The discrepancies revealed by LPF evaluation demonstrate the urgent need for standardised, reproducible evaluation practices in CTG. Our results suggest that without such practices, published performance claims may substantially misrepresent true system capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。