ModernBERT在日文胸部CT报告分类中更高效,训练推理更快,但对真实语境适应性弱。
ModernBERT is More Efficient than Conventional BERT for Chest CT Findings Classification in Japanese Radiology Reports
- 采用相同条件对比三种日文模型,ModernBERT生成更少文本,速度更快
- 内部测试准确率74.7%,优于BERT Base的72.7%,但在外部数据下降明显
- 适合追求效率且数据环境一致的临床场景,需更多真实语料提升鲁棒性
日本医学文本分类面临放射科报告中复杂词汇与语言结构的挑战。本研究在相同条件下对比了BERT Base、JMedRoBERTa和ModernBERT三种日文模型,针对18种胸部CT病灶进行多标签分类。基于CT-RATE-JPN数据集微调后,ModernBERT展现出显著效率优势:生成文本量更少,训练与推理速度更快,内部测试集精确匹配准确率达74.7%,优于BERT Base的72.7%。为评估泛化能力,另构建了243份自然书写日文放射科报告的外部数据集RR-Findings,使用相同标注体系。在域偏移设置下,性能差异显著:BERT Base表现最优,而ModernBERT精确匹配准确率下降最明显。平均精度差异较小,表明其排序能力仍可接受。总体而言,ModernBERT具备高计算效率和良好领域内表现,但对真实语言变异性敏感,提示需更多样化的自然语言训练数据及领域特定校准策略,以提升在异质临床环境中的部署鲁棒性。
原文摘要 · Abstract (English)
Japanese language models for medical text classification face challenges with complex vocabulary and linguistic structures in radiology reports. This study compared three Japanese models--BERT Base, JMedRoBERTa, and ModernBERT--for multi-label classification of 18 chest CT findings. Using the CT-RATE-JPN dataset, all models were fine-tuned under identical conditions. ModernBERT showed clear efficiency advantages, producing substantially fewer tokens and achieving faster training and inference than the other models while maintaining comparable performance on the internal test dataset (exact match accuracy: 74.7% vs. 72.7% for BERT Base). To assess generalizability, we additionally constructed RR-Findings, an external dataset of 243 naturally written Japanese radiology reports annotated using the same schema. Under this domain-shifted setting, performance differences became pronounced: BERT Base outperformed both JMedRoBERTa and ModernBERT, whereas ModernBERT showed the largest decline in exact match accuracy. Average precision differences were smaller, indicating that ModernBERT retained reasonable ranking ability despite reduced calibration. Overall, ModernBERT offers substantial computational efficiency and strong in-domain performance but remains sensitive to real-world linguistic variability. These results highlight the need for more diverse natural-language training data and domain-specific calibration strategies to improve robustness when deploying modern transformer models in heterogeneous clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。