arXiv:2409.18319cs.AIcs.CL2024-09被引 1

开发开源模型自动生成标准化肺结节报告,准确率超97%且无幻觉。

Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports

  • 用动态模板约束解码,从自由文本生成结构化报告。
  • 跨机构测试F1达97%,优于GPT-4o 17.19%。
  • 支持结节检索与统计分析,适合医学研究与本地部署。

现有用于生成完整结构化报告的大型语言模型存在格式错误、内容幻觉及数据上传隐私泄露问题。本文旨在开发一个开源、高精度的LLM,从不同机构的自由文本报告中生成标准化的LCS(Lung Cancer Screening)报告,并展示其在自动统计分析和个体肺结节检索中的应用价值。研究获得伦理委员会批准,回顾性纳入来自两家机构的5,442份去标识化的LDCT LCS放射科报告。通过标注500对自由文本与结构化报告,构建了两个评估数据集,并建立2021年1月至2023年12月的大规模连续数据集。两名放射科医生制定了包含27个肺结节特征的标准记录模板。设计动态模板约束解码方法,增强现有LLM生成结构化报告的能力。利用连续生成的结构化报告,实现描述性统计分析自动化与结节检索原型系统。最优模型在跨机构数据集上表现优异,F1分数约97%,无格式错误或内容幻觉。该方法相较最佳开源LLM提升高达10.42%,优于GPT-4o 17.19%。自动生成的统计分布与既往关于密度、位置、大小、稳定性及Lung-RADS的研究结论一致。基于结构化报告的检索系统支持灵活的结节级搜索与复杂统计分析。所开发软件已公开,支持本地部署与后续研究。

原文摘要 · Abstract (English)

Current LLMs for creating fully-structured reports face the challenges of formatting errors, content hallucinations, and privacy leakage issues when uploading data to external servers.We aim to develop an open-source, accurate LLM for creating fully-structured and standardized LCS reports from varying free-text reports across institutions and demonstrate its utility in automatic statistical analysis and individual lung nodule retrieval. With IRB approvals, our retrospective study included 5,442 de-identified LDCT LCS radiology reports from two institutions. We constructed two evaluation datasets by labeling 500 pairs of free-text and fully-structured radiology reports and one large-scale consecutive dataset from January 2021 to December 2023. Two radiologists created a standardized template for recording 27 lung nodule features on LCS. We designed a dynamic-template-constrained decoding method to enhance existing LLMs for creating fully-structured reports from free-text radiology reports. Using consecutive structured reports, we automated descriptive statistical analyses and a nodule retrieval prototype. Our best LLM for creating fully-structured reports achieved high performance on cross-institutional datasets with an F1 score of about 97%, with neither formatting errors nor content hallucinations. Our method consistently improved the best open-source LLMs by up to 10.42%, and outperformed GPT-4o by 17.19%. The automatically derived statistical distributions were consistent with prior findings regarding attenuation, location, size, stability, and Lung-RADS. The retrieval system with structured reports allowed flexible nodule-level search and complex statistical analysis. Our developed software is publicly available for local deployment and further research.

医学AI结构化报告肺结节大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。