arXiv:2607.09932cs.CLcs.AI2026-07

评估并提升大模型生成临床试验摘要的准确性,避免虚构信息。

Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

论文配图:Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences
图 1 · 摘自论文原文
  • 构建多受众评估框架,用200个真实试验测试大模型摘要。
  • 发现三款主流模型平均有1.55/3的虚假陈述,主要问题为无依据断言。
  • 引入知识图谱检索系统,显著降低错误率,尤其适合医疗领域应用。

大语言模型在医疗领域用于生成临床试验摘要,但其易产生幻觉的问题带来高风险。本研究提出一个基准评估框架,针对医护人员、患者和支付方三类受众,基于从ClinicalTrials.gov数据库提取的200个分层试验,采用特定提示模板与六维忠实度标注体系进行评估。对GPT-4o、Claude Sonnet 4.6和Gemini 2.5 Flash共1800份摘要进行评分,使用交叉编码器自然语言推断(NLI)模型。结果显示,所有模型均以‘无支持声明’为主要失败模式,平均得分1.55/3。开发一种基于知识图谱的检索系统,经验证可显著提升NLI忠诚度得分(蕴含+0.0125,忠实度+0.0130,p < 0.0001)。不同模型改进路径各异:GPT-4o通过减少矛盾改善,而Claude Sonnet 4.6与Gemini 2.5 Flash则通过增强蕴含关系提升性能。

原文摘要 · Abstract (English)

Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to hallucinate poses significant risks in this high-stakes context. This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences. The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templates and a six-dimension faithfulness annotation schema. Baseline measurements were established for GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash across 1,800 generated summaries scored using a cross-encoder natural language inference (NLI) model. Unsupported Claims was identified as the dominant failure mode across all three models, with a mean annotation score of 1.55 out of three. A knowledge-graph-augmented retrieval system was developed and evaluated against the baseline, producing statistically significant improvements in NLI-based faithfulness scores (entailment +0.0125, faithfulness +0.0130, p < 0.0001). Improvement pathways were model-dependent, with GPT-4o improving primarily through contradiction reduction while Claude Sonnet 4.6 and Gemini 2.5 Flash improved through increased entailment.

大模型临床试验忠实度医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。