arXiv:2511.02839cs.HCcs.AI2025-11

GPT-4o可自动识别放射科住院医师报告中的常见错误并提供有效反馈。

Evaluating Generative AI as an Educational Tool for Radiology Resident Report Drafting

  • 用GPT-4o分析真实临床报告,识别三类关键错误。
  • 对三类错误的判断与专家共识匹配度超90%,反馈被评有用率达83%以上。
  • 适合希望提升报告能力的住院医师和资源紧张的教育团队使用。

放射科住院医师需及时、个性化的反馈以提升影像分析与报告能力,但临床负荷常限制导师指导。本研究评估了一种符合HIPAA标准的GPT-4o系统,在多院区美国医疗系统中对住院医师撰写的乳腺影像报告提供自动化反馈。分析了5,000对住院医师-主治医师报告对,通过提示引导GPT-4o识别常见错误。100对报告进行读者研究:四位主治医师与四位住院医师独立判断预设错误类型,并评估GPT-4o反馈是否有助。采用百分比匹配评估一致性,用Krippendorff's alpha衡量读者间可靠性。结果显示,三类常见错误为:(1)关键发现遗漏或添加,(2)技术描述使用或遗漏错误,(3)结论与发现不一致。GPT-4o与主治医师共识匹配率分别为90.5%、78.3%、90.4%;读者间可靠性中等(α=0.767、0.595、0.567),以GPT替代一名人类读者对一致性影响极小(Δ=-0.004至0.002)。反馈在多数情况下被认为有帮助:89.8%、83.0%、92.0%。讨论指出,GPT-4o能可靠识别教育性错误,具备作为可扩展教学工具的潜力。

原文摘要 · Abstract (English)

Objective: Radiology residents require timely, personalized feedback to develop accurate image analysis and reporting skills. Increasing clinical workload often limits attendings' ability to provide guidance. This study evaluates a HIPAA-compliant GPT-4o system that delivers automated feedback on breast imaging reports drafted by residents in real clinical settings. Methods: We analyzed 5,000 resident-attending report pairs from routine practice at a multi-site U.S. health system. GPT-4o was prompted with clinical instructions to identify common errors and provide feedback. A reader study using 100 report pairs was conducted. Four attending radiologists and four residents independently reviewed each pair, determined whether predefined error types were present, and rated GPT-4o's feedback as helpful or not. Agreement between GPT and readers was assessed using percent match. Inter-reader reliability was measured with Krippendorff's alpha. Educational value was measured as the proportion of cases rated helpful. Results: Three common error types were identified: (1) omission or addition of key findings, (2) incorrect use or omission of technical descriptors, and (3) final assessment inconsistent with findings. GPT-4o showed strong agreement with attending consensus: 90.5%, 78.3%, and 90.4% across error types. Inter-reader reliability showed moderate variability (α = 0.767, 0.595, 0.567), and replacing a human reader with GPT-4o did not significantly affect agreement (Δ = -0.004 to 0.002). GPT's feedback was rated helpful in most cases: 89.8%, 83.0%, and 92.0%. Discussion: ChatGPT-4o can reliably identify key educational errors. It may serve as a scalable tool to support radiology education.

AI辅助教学放射科教育生成式AI医学报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。