研究生成数据中的幻觉如何破坏关系抽取效果,发现相关幻觉致召回率下降近40%。
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
- 区分幻觉相关性,发现仅相关幻觉显著影响模型性能
- 幻觉导致召回率下降19.1%至39.2%,相关幻觉影响远大于无关幻觉
- 提出检测方法,可准确识别幻觉文本,助力高质量数据筛选
关系抽取对构建知识图谱至关重要,大规模高质量数据集是训练与评估模型的基础。生成式数据增强(GDA)常用于扩充数据集,但往往引入幻觉(如虚假事实),其对关系抽取的影响尚未充分研究。本文从文档和句子层面实证分析幻觉的影响,发现幻觉显著削弱模型提取关系的能力,召回率下降19.1%至39.2%。我们识别出相关幻觉严重损害模型表现,而无关幻觉影响微乎其微。此外,我们提出幻觉检测方法,能有效区分‘幻觉’与‘纯净’文本,F1分数分别达83.8%和92.2%。该方法不仅可用于清除幻觉,还可估计数据集中幻觉占比,对选择高质量数据具有关键意义。总体而言,本工作证实了相关幻觉对关系抽取模型有效性具有深远影响。
原文摘要 · Abstract (English)
Relation extraction is crucial for constructing knowledge graphs, with large high-quality datasets serving as the foundation for training, fine-tuning, and evaluating models. Generative data augmentation (GDA) is a common approach to expand such datasets. However, this approach often introduces hallucinations, such as spurious facts, whose impact on relation extraction remains underexplored. In this paper, we examine the effects of hallucinations on the performance of relation extraction on the document and sentence levels. Our empirical study reveals that hallucinations considerably compromise the ability of models to extract relations from text, with recall reductions between 19.1% and 39.2%. We identify that relevant hallucinations impair the model's performance, while irrelevant hallucinations have a minimal impact. Additionally, we develop methods for the detection of hallucinations to improve data quality and model performance. Our approaches successfully classify texts as either 'hallucinated' or 'clean,' achieving high F1-scores of 83.8% and 92.2%. These methods not only assist in removing hallucinations but also help in estimating their prevalence within datasets, which is crucial for selecting high-quality data. Overall, our work confirms the profound impact of relevant hallucinations on the effectiveness of relation extraction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。