开源多语言放射科报告数据集,助力跨语言医疗AI研究
PARROT: An Open Multilingual Radiology Reports Dataset
- 收集76位医生在21国撰写的2658份虚构放射科报告,覆盖13种语言
- 报告涵盖CT、MRI等多模态,胸部、腹部等常见部位占近七成
- 首次测试人类与AI生成报告的辨识能力,医生准确率仅略超随机
目的:构建并验证PARROT(多语种标注放射科报告开放测试数据集),一个大规模、多中心、公开可访问的虚构放射科报告数据集,用于测试放射科自然语言处理应用。方法:2024年5月至9月,邀请放射科医生按标准报告流程提交虚构报告。每位贡献者至少提供20份报告,并附带解剖部位、影像模态、临床背景等元数据,非英语报告需提供英文翻译。所有报告均标注ICD-10编码。开展人类与AI报告区分研究,共154名参与者(放射科医生、医疗从业者及非医疗人员)评估报告是否为人工撰写或由AI生成。结果:数据集包含来自76名作者、21个国家、13种语言的2,658份报告,涵盖多种影像模态(CT: 36.1%,MRI: 22.8%,X线: 19.0%,超声: 16.8%)和解剖部位,其中胸部(19.9%)、腹部(18.6%)、头部(17.3%)和骨盆(14.1%)最常见。在区分实验中,参与者平均准确率为53.9%(95%置信区间:50.7%-57.1%),放射科医生表现显著优于其他群体(56.9%,95%置信区间:53.3%-60.6%,p<0.05)。结论:PARROT是目前最大的开放多语言放射科报告数据集,可在无隐私风险下推动跨语言、跨地域、跨临床场景的自然语言处理技术开发与验证。
原文摘要 · Abstract (English)
Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural language processing applications in radiology. Materials and Methods: From May to September 2024, radiologists were invited to contribute fictional radiology reports following their standard reporting practices. Contributors provided at least 20 reports with associated metadata including anatomical region, imaging modality, clinical context, and for non-English reports, English translations. All reports were assigned ICD-10 codes. A human vs. AI report differentiation study was conducted with 154 participants (radiologists, healthcare professionals, and non-healthcare professionals) assessing whether reports were human-authored or AI-generated. Results: The dataset comprises 2,658 radiology reports from 76 authors across 21 countries and 13 languages. Reports cover multiple imaging modalities (CT: 36.1%, MRI: 22.8%, radiography: 19.0%, ultrasound: 16.8%) and anatomical regions, with chest (19.9%), abdomen (18.6%), head (17.3%), and pelvis (14.1%) being most prevalent. In the differentiation study, participants achieved 53.9% accuracy (95% CI: 50.7%-57.1%) in distinguishing between human and AI-generated reports, with radiologists performing significantly better (56.9%, 95% CI: 53.3%-60.6%, p<0.05) than other groups. Conclusion: PARROT represents the largest open multilingual radiology report dataset, enabling development and validation of natural language processing applications across linguistic, geographic, and clinical boundaries without privacy constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。