提出可分解的实体关系幻觉评估框架,提升摘要忠实性检测精度。
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization

- 基于依存句法与实体对齐的改进关系抽取方法
- 实测多模型幻觉率差异显著,最高达37.6%
- 适合关注摘要可靠性与模型可解释性的研究者
抽象式文本摘要系统常在生成流畅内容时虚构或扭曲实体与事件间的关系,导致信息失真。本文提出一种更精准、具依据的关系级幻觉评估框架。通过引入依赖感知的关系抽取算法,结合词形还原、命名实体锚定主语、被动语态恢复、否定词敏感动词建模、报告动词过滤、名词关系回退、从句传播及系统性去重等技术,提升了抽取关系三元组的结构保真度,减少评估中的误匹配。同时提出归一化RHI指标,实现跨数据集和模型的尺度不变比较。新指标将幻觉分解为可解释成分,并整合为归一化的关系忠实度分数。在多个前沿摘要模型上的实证表明,该方法能获得更稳定且区分度更高的幻觉测量结果。本框架推动了自动化关系级忠实性评估的发展,支持面向连贯性与幻觉敏感性的模型分析。
原文摘要 · Abstract (English)
Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of generated summaries, particularly in high-stakes domains. In this work, we present a refined and grounded framework for evaluating relation hallucination in abstractive summarization. We present the empirical Relation Hallucination Index (RHI) by introducing a dependency-aware relation extraction algorithm that incorporates lemmatization-based normalization, named entity grounded subject resolution, passive agent recovery, negation-aware verb modeling, reporting verb filtering, nominal relation fallback, clausal propagation, and systematic deduplication. These enhancements improve the structural fidelity of extracted relation triples and reduce spurious matches during evaluation. In addition, we introduce a normalized formulation of RHI to ensure scale-invariant comparison between datasets and models. The revised metric decomposes hallucination into interpretable components, aggregates relation hallucination metric into a normalized relation faithfulness score. Extensive evaluation across multiple state-of-the-art summarization models demonstrates that the grounded extraction process yields more stable and discriminative hallucination measurements. The proposed framework advances automated relation-level faithfulness evaluation and supports coherence-aware, hallucination-sensitive model analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。