arXiv:2410.21969cs.CV2024-10NeurIPS被引 9

构建统一基准框架,让胸部X光医学视觉语言模型可比可测。

BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

  • 整合九个数据集四类任务,统一评估标准。
  • 验证九个主流模型,部分旧方法性能反超新模型。
  • 适合研究者对比模型、评估泛化能力。

医学视觉语言预训练(MedVLP)在从配对与非配对医学影像和报告中学习通用可迁移表征方面展现出潜力,能为下游任务提供有用特征,并通过少量样本实现新场景的模型适配。然而现有方法在数据集、预处理和微调实现上差异显著,缺乏统一、标准化且全面的评估基准,难以衡量其在多种临床相关任务中的泛化能力。为此,我们提出BenchX,一个统一的基准框架,支持使用公开胸部X光数据集对不同MedVLP方法进行直接比较与系统分析。BenchX包含三个组件:1)涵盖九个数据集和四类医学任务的综合性数据集;2)标准化数据预处理、训练测试划分及参数选择的基准套件;3)统一的微调协议,可兼容异构的MedVLP方法,分别用于分类、分割和报告生成任务。基于BenchX,我们为九个先进MedVLP方法建立了基线,并发现部分早期方法经优化后性能超越近期模型,促使重新审视之前研究的发展与结论。代码已开源。

原文摘要 · Abstract (English)

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX.

医学视觉多模态基准测试胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。