用多任务模型提升肺结核X光片的诊断与分析能力。
PaliGemma-CXR: A Multi-task Multimodal Model for TB Chest X-ray Interpretation
- 基于PaliGemma构建多任务模型,统一处理诊断、检测、分割等
- 诊断准确率达90.32%,闭合式问答正确率98.95%
- 适合医疗影像自动化、多任务模型研究者使用
结核病(TB)是全球性的重大公共卫生挑战。胸部X光片是筛查结核的标准手段,但许多国家面临放射科医生严重短缺的问题。机器学习可提供替代方案,实现疾病诊断和报告生成的自动化。然而,传统方法依赖任务专用模型,无法利用任务间的关联性。构建能执行多项任务的多任务模型面临多重挑战,包括多模态数据稀缺、数据集不平衡及负迁移问题。为此,我们提出PaliGemma-CXR,一个可同时完成结核病诊断、目标检测、图像分割、报告生成和视觉问答(VQA)的多任务多模态模型。基于包含结核病诊断标签和分割掩码的胸片数据集,我们构建了支持多任务的多模态数据集。通过在该数据集上微调PaliGemma,并采用逆数据集大小比例采样策略,所有任务均取得显著成果:诊断准确率为90.32%,闭合式VQA准确率达98.95%,报告生成的BLEU得分为41.3,目标检测和分割的mAP分别为19.4和16.0。这些结果表明,PaliGemma-CXR有效利用了多任务间的内在关联,提升了整体性能。
原文摘要 · Abstract (English)
Tuberculosis (TB) is a infectious global health challenge. Chest X-rays are a standard method for TB screening, yet many countries face a critical shortage of radiologists capable of interpreting these images. Machine learning offers an alternative, as it can automate tasks such as disease diagnosis, and report generation. However, traditional approaches rely on task-specific models, which cannot utilize the interdependence between tasks. Building a multi-task model capable of performing multiple tasks poses additional challenges such as scarcity of multimodal data, dataset imbalance, and negative transfer. To address these challenges, we propose PaliGemma-CXR, a multi-task multimodal model capable of performing TB diagnosis, object detection, segmentation, report generation, and VQA. Starting with a dataset of chest X-ray images annotated with TB diagnosis labels and segmentation masks, we curated a multimodal dataset to support additional tasks. By finetuning PaliGemma on this dataset and sampling data using ratios of the inverse of the size of task datasets, we achieved the following results across all tasks: 90.32% accuracy on TB diagnosis and 98.95% on close-ended VQA, 41.3 BLEU score on report generation, and a mAP of 19.4 and 16.0 on object detection and segmentation, respectively. These results demonstrate that PaliGemma-CXR effectively leverages the interdependence between multiple image interpretation tasks to enhance performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。