构建脊柱疾病分析专用多模态大模型评测基准
SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
- 基于4万张脊柱影像构建6.5万组问答对,覆盖11类疾病
- 12个主流多模态模型在该基准上表现普遍不佳
- 专为真实临床场景设计,适合医学AI研究者使用
随着多模态大语言模型(MLLMs)在医疗领域的广泛应用,对其在不同医疗领域表现的全面评估变得至关重要。然而,现有基准主要评估通用医疗任务,未能充分反映模型在依赖视觉输入的精细领域如脊柱诊断中的表现。为此,我们提出了SpineBench,一个针对脊柱领域细粒度分析与评估的视觉问答(VQA)基准。SpineBench包含来自40,263张脊柱图像的64,878组问答对,涵盖11种脊柱疾病,通过两种关键临床任务——脊柱疾病诊断与脊柱病灶定位——进行多选题测评。该基准通过整合和标准化开源脊柱疾病数据集的图像-标签对,并为每组问答生成视觉相似但不相同的干扰项(hard negative),模拟真实临床中难以区分的挑战场景。我们在SpineBench上评估了12个领先MLLMs,结果显示这些模型在脊柱任务中表现普遍较差,揭示了当前MLLM在脊柱领域的局限性,为未来脊柱医学应用的改进提供了方向。SpineBench已公开发布于 https://zhangchenghanyu.github.io/SpineBench.github.io/。
原文摘要 · Abstract (English)
With the increasing integration of Multimodal Large Language Models (MLLMs) into the medical field, comprehensive evaluation of their performance in various medical domains becomes critical. However, existing benchmarks primarily assess general medical tasks, inadequately capturing performance in nuanced areas like the spine, which relies heavily on visual input. To address this, we introduce SpineBench, a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of MLLMs in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical clinical tasks: spinal disease diagnosis and spinal lesion localization, both in multiple-choice format. SpineBench is built by integrating and standardizing image-label pairs from open-source spinal disease datasets, and samples challenging hard negative options for each VQA pair based on visual similarity (similar but not the same disease), simulating real-world challenging scenarios. We evaluate 12 leading MLLMs on SpineBench. The results reveal that these models exhibit poor performance in spinal tasks, highlighting limitations of current MLLM in the spine domain and guiding future improvements in spinal medicine applications. SpineBench is publicly available at https://zhangchenghanyu.github.io/SpineBench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。