arXiv:2501.16309physics.med-phcs.AI2025-01被引 8

用大模型自动总结放疗CT检查单,准确率超98%。

Evaluating The Performance of Using Large Language Models to Automate Summarization of CT Simulation Orders in Radiation Oncology

  • 用Llama 3.1 405B模型结合定制提示词生成摘要。
  • 98%的摘要与人工基准一致,格式更统一、可读性更强。
  • 适用于各类放疗模态和病灶部位,适合临床流程优化。

目的:本研究旨在利用大语言模型(LLM)自动化生成放疗CT模拟订单的摘要,并评估其性能。方法:从本机构Aria数据库收集了607份患者CT模拟订单,采用本地部署的Llama 3.1 405B模型通过API接口提取关键词并生成摘要。订单按治疗模态和疾病部位分为七类,每类均与治疗师协作设计定制化指令提示,引导模型生成摘要。每份摘要的真值由治疗师仔细审阅后人工构建并验证。通过对比真值评估模型生成摘要的准确性。结果:约98%的模型生成摘要在准确性上与人工基准一致。评估显示,模型生成摘要在格式一致性及可读性方面优于治疗师手动撰写版本,且在所有模态和病灶部位中表现稳定。结论:Llama 3.1 405B模型在提取关键词和生成摘要方面表现出高精度与一致性,表明大模型在此任务中具有巨大潜力,可减轻治疗师工作负担,提升流程效率。

原文摘要 · Abstract (English)

Purpose: This study aims to use a large language model (LLM) to automate the generation of summaries from the CT simulation orders and evaluate its performance. Materials and Methods: A total of 607 CT simulation orders for patients were collected from the Aria database at our institution. A locally hosted Llama 3.1 405B model, accessed via the Application Programming Interface (API) service, was used to extract keywords from the CT simulation orders and generate summaries. The downloaded CT simulation orders were categorized into seven groups based on treatment modalities and disease sites. For each group, a customized instruction prompt was developed collaboratively with therapists to guide the Llama 3.1 405B model in generating summaries. The ground truth for the corresponding summaries was manually derived by carefully reviewing each CT simulation order and subsequently verified by therapists. The accuracy of the LLM-generated summaries was evaluated by therapists using the verified ground truth as a reference. Results: About 98% of the LLM-generated summaries aligned with the manually generated ground truth in terms of accuracy. Our evaluations showed an improved consistency in format and enhanced readability of the LLM-generated summaries compared to the corresponding therapists-generated summaries. This automated approach demonstrated a consistent performance across all groups, regardless of modality or disease site. Conclusions: This study demonstrated the high precision and consistency of the Llama 3.1 405B model in extracting keywords and summarizing CT simulation orders, suggesting that LLMs have great potential to help with this task, reduce the workload of therapists and improve workflow efficiency.

大模型医疗自动化放疗摘要生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。