arXiv:2510.03160cs.CVcs.AI2025-10被引 3

首个面向脊柱影像的多模态水平级基准,助力精准诊断。

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus

  • 构建涵盖45万条指令的SpineMed-450k数据集,支持多模态脊柱影像水平级推理
  • 在脊柱疾病诊断中,微调模型在定位、病灶评估等任务上显著优于主流模型
  • 专为临床医生设计,适合医疗AI研发与脊柱疾病辅助诊断场景

脊柱疾病影响全球6.19亿人,是致残主因之一,但现有AI辅助诊断受限于缺乏水平级、多模态数据集。临床决策需跨X-ray、CT、MRI在特定椎体水平进行复杂推理,但缺乏可追溯的临床指导数据与标准化基准。为此,我们提出SpineMed生态,包含由脊柱外科医生共同设计的SpineMed-450k数据集——首个专为椎体水平推理构建的大规模多模态数据集,含超45万条指令实例,数据源自教材、指南、开源数据及约1,000例去标识化医院病例,采用两阶段大语言模型生成流程(初稿+修订)确保高质量与可追溯性,适用于问答、多轮咨询与报告生成。同时推出SpineBench,一个以临床相关性为核心的评估框架,涵盖水平识别、病理评估与手术规划等维度。对多个先进视觉-语言模型的全面评估显示其在细粒度水平推理中存在系统性缺陷;而基于SpineMed-450k微调的模型在所有任务中均实现一致且显著提升,临床医生评估也证实其输出具备诊断清晰性与实用性。

原文摘要 · Abstract (English)

Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets. Clinical decision-making for spine disorders requires sophisticated reasoning across X-ray, CT, and MRI at specific vertebral levels. However, progress has been constrained by the absence of traceable, clinically-grounded instruction data and standardized, spine-specific benchmarks. To address this, we introduce SpineMed, an ecosystem co-designed with practicing spine surgeons. It features SpineMed-450k, the first large-scale dataset explicitly designed for vertebral-level reasoning across imaging modalities with over 450,000 instruction instances, and SpineBench, a clinically-grounded evaluation framework. SpineMed-450k is curated from diverse sources, including textbooks, guidelines, open datasets, and ~1,000 de-identified hospital cases, using a clinician-in-the-loop pipeline with a two-stage LLM generation method (draft and revision) to ensure high-quality, traceable data for question-answering, multi-turn consultations, and report generation. SpineBench evaluates models on clinically salient axes, including level identification, pathology assessment, and surgical planning. Our comprehensive evaluation of several recently advanced large vision-language models (LVLMs) on SpineBench reveals systematic weaknesses in fine-grained, level-specific reasoning. In contrast, our model fine-tuned on SpineMed-450k demonstrates consistent and significant improvements across all tasks. Clinician assessments confirm the diagnostic clarity and practical utility of our model's outputs.

医学影像多模态脊柱诊断数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。