对比5种AI模型在急诊胸片报告生成中的表现,发现AIRead综合表现最佳。
Comparative Evaluation of Generative AI Models for Chest Radiograph Report Generation in the Emergency Department
- 用放射科医生报告作基准,盲评5个视觉语言模型的报告质量。
- AIRead在临床可用性、错误率和幻觉控制上优于其他模型,敏感性也更稳定。
- 适合临床评估医疗AI模型性能,尤其关注报告准确性和可靠性的人参考。
目的:对比开源或商用医学图像专用视觉语言模型(VLM)与真实放射科医生撰写的报告。方法:回顾性研究纳入2022年1月至2025年4月期间急诊科就诊且当日完成胸部X光(CXR)和CT检查的成年患者。随机呈现五种VLM(AIRead、Lingshu、MAIRA-2、MedGemma、MedVersa)及放射科医生报告,由三位胸腔放射科医师采用四个标准(RADPEER、临床可接受性、幻觉、语言清晰度)进行盲评。使用广义线性混合模型评估性能,以放射科医生报告为参考。还以CT为金标准进行发现级别分析。结果:共纳入478例患者(中位年龄67岁,四分位间距50–78;男性282例,占59.0%)。AIRead的RADPEER 3b率最低(5.3% [76/1434],而放射科医生为13.9% [200/1434];P<.001),其余模型差异率更高(16.8%-43.0%;P<.05)。临床可接受性最高的是AIRead(84.5% [1212/1434],放射科医生74.3% [1065/1434];P<.001),其他模型较差(41.1%-71.4%;P<.05)。AIRead幻觉罕见,与放射科医生相当(0.3% [4/1425] 对比 0.1% [1/1425];P=.21),但其他模型幻觉频发(5.4%-17.4%;P<.05)。语言清晰度方面,AIRead(82.9% [1189/1434])、Lingshu(88.0% [1262/1434])和MedVersa(88.4% [1268/1434])高于放射科医生(78.1% [1120/1434];P<.05)。不同模型对常见发现的敏感性差异显著:AIRead(15.5%-86.7%)、Lingshu(2.4%-86.7%)、MAIRA-2(6.0%-72.0%)、MedGemma(4.8%-76.7%)、MedVersa(20.2%-69.3%)。结论:用于胸片报告生成的医学视觉语言模型在报告质量与诊断性能上表现不一。
原文摘要 · Abstract (English)
Purpose: To benchmark open-source or commercial medical image-specific VLMs against real-world radiologist-written reports. Methods: This retrospective study included adult patients who presented to the emergency department between January 2022 and April 2025 and underwent same-day CXR and CT for febrile or respiratory symptoms. Reports from five VLMs (AIRead, Lingshu, MAIRA-2, MedGemma, and MedVersa) and radiologist-written reports were randomly presented and blindly evaluated by three thoracic radiologists using four criteria: RADPEER, clinical acceptability, hallucination, and language clarity. Comparative performance was assessed using generalized linear mixed models, with radiologist-written reports treated as the reference. Finding-level analyses were also performed with CT as the reference. Results: A total of 478 patients (median age, 67 years [interquartile range, 50-78]; 282 men [59.0%]) were included. AIRead demonstrated the lowest RADPEER 3b rate (5.3% [76/1434] vs. radiologists 13.9% [200/1434]; P<.001), whereas other VLMs showed higher disagreement rates (16.8-43.0%; P<.05). Clinical acceptability was the highest with AIRead (84.5% [1212/1434] vs. radiologists 74.3% [1065/1434]; P<.001), while other VLMs performed worse (41.1-71.4%; P<.05). Hallucinations were rare with AIRead, comparable to radiologists (0.3% [4/1425]) vs. 0.1% [1/1425]; P=.21), but frequent with other models (5.4-17.4%; P<.05). Language clarity was higher with AIRead (82.9% [1189/1434]), Lingshu (88.0% [1262/1434]), and MedVersa (88.4% [1268/1434]) compared with radiologists (78.1% [1120/1434]; P<.05). Sensitivity varied substantially across VLMs for the common findings: AIRead, 15.5-86.7%; Lingshu, 2.4-86.7%; MAIRA-2, 6.0-72.0%; MedGemma, 4.8-76.7%; and MedVersa, 20.2-69.3%. Conclusion: Medical VLMs for CXR report generation exhibited variable performance in report quality and diagnostic measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。