arXiv:2608.26154cs.CLcs.AI2026-08

评估大模型生成的癌症患者摘要质量,提升医疗文本安全与准确

Evaluating AI Generated Summaries for Cancer Patients

论文配图:Evaluating AI Generated Summaries for Cancer Patients
图 1 · 摘自论文原文
  • 用临床专家和大模型双轨评估摘要质量
  • 发现摘要存在遗漏和轻微错误,影响临床可靠性
  • 通过迭代优化提示词和安全机制提升生成效果

大型语言模型(LLMs)正被越来越多地用于数字健康平台,以生成复杂医学数据的摘要。尽管这些模型能提升患者参与度和沟通效率,但在临床环境中仍存在准确性、忠实性与安全性方面的担忧。本研究在癌症患者护理应用中,采用双评估框架评估AI生成摘要的质量。由肿瘤科医生及面向患者的护理人员组成的领域专家对摘要的准确性、临床相关性和可读性进行基准评估。同时,我们使用大模型作为评估者(LLM-as-a-judge)。分析发现生成摘要存在偶尔遗漏和细微不准确等问题,这些问题被系统性识别,并用于迭代优化提示设计、内容锚定和安全防护机制。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfulness, and safety in clinical contexts. In this study, we evaluate AI-generated summaries within a cancer patient care application using a dual assessment framework. Human domain experts, including oncology clinicians and patient-facing care staff, provided ground-truth evaluations of summary quality along dimensions of accuracy, clinical relevance, and readability. In parallel, we employed LLMs serving as evaluators (LLM-as-a-judge). Some limitations were identified in the generated summaries e.g., occasional omissions and minor inaccuracies. These were systematically analyzed and used to iteratively improve prompt design, grounding, and safety guardrails.

医疗AI摘要生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。