arXiv:2503.08890cs.CL2025-03中稿 · Journal of Biomedi…被引 4

针对医学通俗摘要的幻觉问题,提出新评估方法PlainQAFact。

PlainQAFact: Retrieval-augmented Factual Consistency Evaluation Metric for Biomedical Plain Language Summarization

  • 基于检索增强的问答机制,区分句子类型后评估事实一致性。
  • 在复杂解释句上表现优于现有方法,相关性提升12.3%。
  • 适合医学信息传播、AI生成内容安全评估等场景使用。

大语言模型在医疗领域生成幻觉内容,对公众健康决策构成风险。现有自动事实一致性评估方法(如蕴含和问答类)因通俗语言摘要中存在扩展解释现象(如定义、背景、例证等),难以准确评估。为此,我们提出PlainQAFact,一种基于细粒度人工标注数据集PlainFact训练的评估指标,用于评估源简化句与扩展解释句的事实一致性。该方法先分类句子类型,再采用检索增强的问答评分策略。实验证明,现有指标在通俗摘要中表现不佳,尤其在扩展解释句上,而PlainQAFact在所有测试场景中均显著优于基准方法。进一步分析表明,其在不同外部知识源、答案提取策略、答案重叠度量和文档粒度下均具鲁棒性,有效提升整体评估能力。本工作为生物医学通俗语言摘要中的扩展解释提供了一种面向句子、检索增强的评估工具,既建立新基准,也推动医疗领域可信赖、安全的通俗传播。

原文摘要 · Abstract (English)

Hallucinated outputs from large language models (LLMs) pose risks in the medical domain, especially for lay audiences making health-related decisions. Existing automatic factual consistency evaluation methods, such as entailment- and question-answering (QA) -based, struggle with plain language summarization (PLS) due to elaborative explanation phenomenon, which introduces external content (e.g., definitions, background, examples) absent from the scientific abstract to enhance comprehension. To address this, we introduce PlainQAFact, an automatic factual consistency evaluation metric trained on a fine-grained, human-annotated dataset PlainFact, for evaluating factual consistency of both source-simplified and elaborately explained sentences. PlainQAFact first classifies sentence type, then applies a retrieval-augmented QA scoring method. Empirical results show that existing evaluation metrics fail to evaluate the factual consistency in PLS, especially for elaborative explanations, whereas PlainQAFact consistently outperforms them across all evaluation settings. We further analyze PlainQAFact's effectiveness across external knowledge sources, answer extraction strategies, answer overlap measures, and document granularity levels, refining its overall factual consistency assessment. Taken together, our work presents a sentence-aware, retrieval-augmented metric targeted at elaborative explanations in biomedical PLS tasks, providing the community with both a new benchmark and a practical evaluation tool to advance reliable and safe plain language communication in the medical domain. PlainQAFact and PlainFact are available at: https://github.com/zhiwenyou103/PlainQAFact

医学AI事实一致性评估指标通俗摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。