arXiv:2410.07473cs.CL2024-10Transactions of th…被引 12

将生成文本拆解为问答对,精准定位事实错误

Localizing Factual Inconsistencies in Attributable Text Generation

  • 把生成内容分解成最小语义单位的问答对
  • 在3000+样本上实现高人工一致性标注
  • 适合需要细粒度事实校验的研究者使用

生成文本中的幻觉检测日益受到关注,但现有方法难以精确定位错误。本文提出QASemConsistency,一种基于新大卫森语义形式的细粒度事实不一致定位方法。将生成文本分解为最小谓词-论元级命题,表示为简单问答对,并判断每个问答对是否被可信参考文本支持。由于每个问答对对应一个谓词与论元间的单一语义关系,该方法能有效定位未被支持的信息。我们通过众包收集了超过3000个实例的细粒度一致性错误标注,实现了较高的标注者间一致性。该基准涵盖多种可归因文本生成任务。实验表明,QASemConsistency得分与人类判断高度相关。此外,我们还实现了基于监督蕴含模型和大语言模型的自动检测方法。

原文摘要 · Abstract (English)

There has been an increasing interest in detecting hallucinations in model-generated texts, both manually and automatically, at varying levels of granularity. However, most existing methods fail to precisely pinpoint the errors. In this work, we introduce QASemConsistency, a new formalism for localizing factual inconsistencies in attributable text generation, at a fine-grained level. Drawing inspiration from Neo-Davidsonian formal semantics, we propose decomposing the generated text into minimal predicate-argument level propositions, expressed as simple question-answer (QA) pairs, and assess whether each individual QA pair is supported by a trusted reference text. As each QA pair corresponds to a single semantic relation between a predicate and an argument, QASemConsistency effectively localizes the unsupported information. We first demonstrate the effectiveness of the QASemConsistency methodology for human annotation, by collecting crowdsourced annotations of granular consistency errors, while achieving a substantial inter-annotator agreement. This benchmark includes more than 3K instances spanning various tasks of attributable text generation. We also show that QASemConsistency yields factual consistency scores that correlate well with human judgments. Finally, we implement several methods for automatically detecting localized factual inconsistencies, with both supervised entailment models and LLMs.

事实检测生成质量语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。