arXiv:2412.15907cs.CLcs.AI2024-12被引 2

构建日语胸部CT报告数据集并开发高精度病灶分类模型

Development of a Large-scale Dataset of Chest Computed Tomography Reports in Japanese and a High-performance Finding Classification Model

  • 用GPT-4o mini机器翻译生成2.4万份日语CT报告,经放射科医生校验
  • 自研模型CT-BERT-JPN在18项病灶识别中14项F1超0.95,部分优于GPT-4o
  • 为日本医学影像语言模型研究提供高质量数据与评估基准

背景:大语言模型的发展亟需高质量多语言医疗数据。尽管日本在CT扫描仪部署和使用方面全球领先,但缺乏大规模日语放射学数据集,制约了医学影像分析专用语言模型的发展。目标:通过机器翻译构建全面的日语CT报告数据集,并建立用于结构化病灶分类的专用语言模型;同时创建经专家放射科医生严格审核的评估数据集。方法:利用GPT-4o mini将CT-RATE数据集(24,283份报告,来自21,304名患者)翻译成日语。训练集包含22,778份机器翻译报告,验证集包含150份放射科医生修订报告。基于"tohoku-nlp/bert-base-japanese-v3"架构开发了CT-BERT-JPN模型,用于从日语放射学报告中提取18种结构化病灶。结果:翻译质量良好,查找部分BLEU得分为0.731,印象部分为0.690,ROUGE得分范围为0.770–0.876(查找)和0.748–0.857(印象)。与GPT-4o相比,CT-BERT-JPN在18项中的11项表现更优,包括淋巴结肿大(+14.2%)、叶间隔增厚(+10.9%)和肺不张(+7.4%)。模型在14项中F1得分超过0.95,在4项中达到完美分数。结论:本研究建立了可靠的日语CT报告数据集,证明了专用语言模型在结构化病灶分类中的有效性。机器翻译与专家验证相结合的混合方法可在保证高质量的前提下构建大规模医疗数据集。

原文摘要 · Abstract (English)

Background: Recent advances in large language models highlight the need for high-quality multilingual medical datasets. While Japan leads globally in CT scanner deployment and utilization, the lack of large-scale Japanese radiology datasets has hindered the development of specialized language models for medical imaging analysis. Objective: To develop a comprehensive Japanese CT report dataset through machine translation and establish a specialized language model for structured finding classification. Additionally, to create a rigorously validated evaluation dataset through expert radiologist review. Methods: We translated the CT-RATE dataset (24,283 CT reports from 21,304 patients) into Japanese using GPT-4o mini. The training dataset consisted of 22,778 machine-translated reports, while the validation dataset included 150 radiologist-revised reports. We developed CT-BERT-JPN based on "tohoku-nlp/bert-base-japanese-v3" architecture for extracting 18 structured findings from Japanese radiology reports. Results: Translation metrics showed strong performance with BLEU scores of 0.731 and 0.690, and ROUGE scores ranging from 0.770 to 0.876 for Findings and from 0.748 to 0.857 for Impression sections. CT-BERT-JPN demonstrated superior performance compared to GPT-4o in 11 out of 18 conditions, including lymphadenopathy (+14.2%), interlobular septal thickening (+10.9%), and atelectasis (+7.4%). The model maintained F1 scores exceeding 0.95 in 14 out of 18 conditions and achieved perfect scores in four conditions. Conclusions: Our study establishes a robust Japanese CT report dataset and demonstrates the effectiveness of a specialized language model for structured finding classification. The hybrid approach of machine translation and expert validation enables the creation of large-scale medical datasets while maintaining high quality.

医学影像自然语言处理日语数据病灶分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。