用大模型提升健康咨询摘要的准确性,避免医疗信息失真。
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
- 结合句子提取与医学实体识别,增强摘要忠实度。
- 在英、孟加拉语数据集上,各项指标均优于基线,关键信息保留率超80%。
- 适合医疗AI部署,尤其关注摘要真实性的研究者与开发者。
消费者健康咨询(CHQs)的摘要有助于改善医疗沟通,但不忠实的摘要可能歪曲医疗细节,带来严重风险。本文提出一种融合TextRank句段提取与医学命名实体识别的大语言模型框架,以提升医疗文本摘要的忠实度。实验中,在MeQSum(英文)和BanglaCHQ-Summ(孟加拉语)数据集上微调LLaMA-2-7B模型,结果在质量(ROUGE、BERTScore、可读性)和忠实度(SummaC、AlignScore)指标上均有稳定提升,显著优于零样本基线与先前系统。人工评估显示,超过80%的生成摘要完整保留了关键医疗信息。研究强调忠实度是可靠医疗摘要的关键维度,并验证了该方法在医疗场景中安全部署LLMs的潜力。
原文摘要 · Abstract (English)
Summarizing consumer health questions (CHQs) can ease communication in healthcare, but unfaithful summaries that misrepresent medical details pose serious risks. We propose a framework that combines TextRank-based sentence extraction and medical named entity recognition with large language models (LLMs) to enhance faithfulness in medical text summarization. In our experiments, we fine-tuned the LLaMA-2-7B model on the MeQSum (English) and BanglaCHQ-Summ (Bangla) datasets, achieving consistent improvements across quality (ROUGE, BERTScore, readability) and faithfulness (SummaC, AlignScore) metrics, and outperforming zero-shot baselines and prior systems. Human evaluation further shows that over 80\% of generated summaries preserve critical medical information. These results highlight faithfulness as an essential dimension for reliable medical summarization and demonstrate the potential of our approach for safer deployment of LLMs in healthcare contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。