arXiv:2505.18452cs.CL2025-05ACL被引 3

针对医学问答的幻觉问题,提出可适配领域的事实性评估新方法。

MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification

  • 通过条件感知分解将医学回答拆解为有效事实
  • 提取的有效事实数量是现有方法的三倍以上
  • 适合医疗领域事实性评估,尤其对临床对话场景有效

大型语言模型虽能生成流畅回答,但未必准确。在流行的“分解-验证”事实性评估范式中,模型需将生成内容拆分为独立有效命题进行判断。这一评估对医学问答尤为重要,因错误信息可能危及患者。然而,现有系统多针对客观、实体中心、公式化文本(如传记、历史)评测,难以应对医学回答中条件依赖、对话性强、结构多样、主观性高的特点,导致事实分解困难。为此,我们提出 MedScore,一种基于领域适配的事实分解与验证新流程,能将医学回答分解为具有条件感知的可靠事实,并在领域内语料库中验证。该方法提取的有效事实数量可达现有方法的三倍,显著减少幻觉与模糊引用,同时保留事实的条件依赖性。结果表明,事实性评分受分解方法、验证语料库及主干模型影响显著,强调了为可靠评估必须定制每个环节。本方法具备通用性与模块化设计,适用于医学领域事实性评估的持续优化。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) can generate fluent and convincing responses, they are not necessarily correct. This is especially apparent in the popular decompose-then-verify factuality evaluation pipeline, where LLMs evaluate generations by decomposing the generations into individual, valid claims. Factuality evaluation is especially important for medical answers, since incorrect medical information could seriously harm the patient. However, existing factuality systems are a poor match for the medical domain, as they are typically only evaluated on objective, entity-centric, formulaic texts such as biographies and historical topics. This differs from condition-dependent, conversational, hypothetical, sentence-structure diverse, and subjective medical answers, which makes decomposition into valid facts challenging. We propose MedScore, a new pipeline to decompose medical answers into condition-aware valid facts and verify against in-domain corpora. Our method extracts up to three times more valid facts than existing methods, reducing hallucination and vague references, and retaining condition-dependency in facts. The resulting factuality score substantially varies by decomposition method, verification corpus, and used backbone LLM, highlighting the importance of customizing each step for reliable factuality evaluation by using our generalizable and modularized pipeline for domain adaptation.

医学问答事实性评估大模型幻觉领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。