AI可精准生成胸片报告,连导管位置变化都识别得准。
Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting
- 用310万份胸片数据训练多模态模型,同时分析病灶和导管
- 在导管类型、位置、变化上准确率超现有技术,关键错误率仅4.6%
- 真实放射科医生盲评证实,AI报告质量接近人工,适合高负荷场景
AI辅助报告生成有望缓解筛查指南扩展、复杂病例增多和人力短缺带来的放射科工作压力,同时保持诊断准确性。除了描述胸部X光片中的病理发现外,解读导管和线路(L&T)对放射科医生来说既耗时又重复,尤其在高患者量下更显负担。我们提出MAIRA-X,一种经临床验证的多模态AI模型,用于纵向胸部X光片(CXR)报告生成,能同时涵盖临床发现和L&T报告。该模型基于来自梅奥诊所的大规模多中心纵向数据集(310万例研究,共600万张图像,涉及80.6万名患者)进行训练,并在三个独立测试集及公开的MIMIC-CXR数据集上评估,显著优于现有技术水平,在词汇质量、临床正确性以及与L&T相关要素方面表现突出。我们开发了一套全新的、针对L&T的评估指标体系,用于衡量类型、纵向变化和放置位置等属性的准确性。首次开展的回顾性用户评估研究中,九位不同经验水平的放射科医生盲审了600个来自不同受试者的案例。结果显示,原始报告与AI生成报告的关键错误率分别为3.0%和4.6%,可接受语句比例分别为97.8%和97.4%,相较以往用户研究有显著改进且误差更低。结果表明,MAIRA-X能有效辅助放射科医生,特别是在高容量临床环境中。
原文摘要 · Abstract (English)
AI-assisted report generation offers the opportunity to reduce radiologists' workload stemming from expanded screening guidelines, complex cases and workforce shortages, while maintaining diagnostic accuracy. In addition to describing pathological findings in chest X-ray reports, interpreting lines and tubes (L&T) is demanding and repetitive for radiologists, especially with high patient volumes. We introduce MAIRA-X, a clinically evaluated multimodal AI model for longitudinal chest X-ray (CXR) report generation, that encompasses both clinical findings and L&T reporting. Developed using a large-scale, multi-site, longitudinal dataset of 3.1 million studies (comprising 6 million images from 806k patients) from Mayo Clinic, MAIRA-X was evaluated on three holdout datasets and the public MIMIC-CXR dataset, where it significantly improved AI-generated reports over the state of the art on lexical quality, clinical correctness, and L&T-related elements. A novel L&T-specific metrics framework was developed to assess accuracy in reporting attributes such as type, longitudinal change and placement. A first-of-its-kind retrospective user evaluation study was conducted with nine radiologists of varying experience, who blindly reviewed 600 studies from distinct subjects. The user study found comparable rates of critical errors (3.0% for original vs. 4.6% for AI-generated reports) and a similar rate of acceptable sentences (97.8% for original vs. 97.4% for AI-generated reports), marking a significant improvement over prior user studies with larger gaps and higher error rates. Our results suggest that MAIRA-X can effectively assist radiologists, particularly in high-volume clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。