首个对大模型生成文本错误全面标注的数据集,助力精准诊断与改进。
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
- 用提示词引导GPT-2生成候选句,人工标注4.7万句中的1.2万错误
- 构建涵盖24类错误的分类体系,每处错误配关联句与修正建议
- 支持自动检测、分类、解释等诊断任务,适合评估和优化生成质量
为深入理解预训练语言模型在文本生成中的能力并开展诊断性评估,我们提出TGEA——一个针对预训练语言模型(PLM)生成文本的错误标注数据集及基准任务。通过精心选择的提示词引导GPT-2生成候选句子,从中选取47,000句进行人工错误标注,共识别出12,000句存在错误。我们建立了一个涵盖24种错误类型的分类体系,依据错误在语言学和常识层面的性质划分。针对每个错误片段,还标注了与其密切相关的另一片段。每条错误均包含错误位置、关联片段、最小修正、错误类型及成因说明等完整标注。除完整标注数据集外,还详细描述了数据收集流程、统计分析与结果。TGEA是首个对PLM生成文本提供全面标注的数据集,可推动基于模型的文本生成诊断评估。此外,我们以TGEA为基准,设计了一系列自动诊断任务,包括错误检测、类型分类、关联片段识别与错误原因生成,进一步促进对大模型生成文本自动纠错的研究。
原文摘要 · Abstract (English)
In order to deeply understand the capability of pretrained language models in text generation and conduct a diagnostic evaluation, we propose TGEA, an error-annotated dataset with multiple benchmark tasks for text generation from pretrained language models (PLMs). We use carefully selected prompt words to guide GPT-2 to generate candidate sentences, from which we select 47K for error annotation. Crowdsourced workers manually check each of these sentences and detect 12k erroneous sentences. We create an error taxonomy to cover 24 types of errors occurring in these erroneous sentences according to the nature of errors with respect to linguistics and knowledge (eg, common sense). For each erroneous span in PLM-generated sentences, we also detect another span that is closely associated with it. Each error is hence manually labeled with comprehensive annotations, including the span of the error, the associated span, minimal correction to the error, the type of the error, and rationale behind the error. Apart from the fully annotated dataset, we also present a detailed description of the data collection procedure, statistics and analysis of the dataset. This is the first dataset with comprehensive annotations for PLM-generated texts, which facilitates the diagnostic evaluation of PLM-based text generation. Furthermore, we use TGEA as a benchmark dataset and propose a series of automatic diagnosis tasks, including error detection, error type classification, associated span detection, error rationale generation, to further promote future study on the automatic error detection and correction on texts generated by pretrained language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。