arXiv:2505.10472cs.CLcs.AI2025-05被引 2

对比通用与医疗大模型在癌症科普中的表现,发现各有优劣。

Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI

  • 用多方法评估5个通用+3个医疗大模型的表达质量、安全性和可读性。
  • 通用模型语言更流畅,医疗模型信息更易懂但更易出现有害内容。
  • 研究提示需针对性优化模型安全性,适合医疗AI开发者参考。

乳腺癌和宫颈癌的有效沟通仍是重大健康挑战,公众对预防、筛查和治疗的认知存在显著缺口,可能导致诊断延迟和治疗不足。本研究评估了大型语言模型(LLMs)在生成准确、安全且可访问的癌症相关信息方面的能力与局限,以支持患者理解。我们采用混合方法评估框架,对五种通用型和三种医疗专用型LLMs进行了语言质量、安全可信度以及沟通可及性与感染力的评估。评估方法包括量化指标、专家定性评分及基于Welch's ANOVA、Games-Howell和Hedges' g的统计分析。结果表明,通用型LLMs在语言质量和情感感染力方面表现更优,而医疗型LLMs在信息可及性上更具优势。然而,医疗型模型表现出更高的潜在危害、毒性与偏见水平,导致其在安全性和可信度上表现较差。研究揭示了领域知识与安全性的双重矛盾。结论强调需要有意识地进行模型设计,尤其在减少伤害与偏见、提升安全性和感染力方面进行针对性改进。本研究为癌症沟通中LLMs的应用提供了全面评估,为未来开发准确、安全、可访问的数字健康工具提供了关键洞见。

原文摘要 · Abstract (English)

Effective communication about breast and cervical cancers remains a persistent health challenge, with significant gaps in public understanding of cancer prevention, screening, and treatment, potentially leading to delayed diagnoses and inadequate treatments. This study evaluates the capabilities and limitations of Large Language Models (LLMs) in generating accurate, safe, and accessible cancer-related information to support patient understanding. We evaluated five general-purpose and three medical LLMs using a mixed-methods evaluation framework across linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness. Our approach utilized quantitative metrics, qualitative expert ratings, and statistical analysis using Welch's ANOVA, Games-Howell, and Hedges' g. Our results show that general-purpose LLMs produced outputs of higher linguistic quality and affectiveness, while medical LLMs demonstrate greater communication accessibility. However, medical LLMs tend to exhibit higher levels of potential harm, toxicity, and bias, reducing their performance in safety and trustworthiness. Our findings indicate a duality between domain-specific knowledge and safety in health communications. The results highlight the need for intentional model design with targeted improvements, particularly in mitigating harm and bias, and improving safety and affectiveness. This study provides a comprehensive evaluation of LLMs for cancer communication, offering critical insights for improving AI-generated health content and informing future development of accurate, safe, and accessible digital health tools.

大模型癌症科普AI医疗安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。