构建放射科空间解剖推理测评数据集,助力医疗视觉语言模型精准理解人体结构关系。
SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models
- 人工设计300组影像-问题对,覆盖多模态医学影像与解剖定位任务
- 涵盖腹部、胸部等五大部位,支持空间关系与器官定位的精细化评估
- 提供标准化评测流程,适合模型开发、故障分析及临床部署前验证
视觉语言模型在医学影像领域日益受到关注,但现有基准测试多聚焦疾病分类或报告生成,缺乏对放射科所需空间与解剖推理能力的评估。我们构建了面向临床放射学的空间感知与解剖推理基准(SPARC-Rad),这是一个人工标注的多模态基准数据集和评估流程,用于评估放射科视觉语言模型在解剖识别、定位、侧别判断、区域认知、器械识别及结构间空间关系等方面的能力。数据集包含300个来自癌症影像档案库(TCIA)健康对照研究的图像-问题对,涵盖腹部、胸部、乳腺、神经及肌肉骨骼五大类别,以及CT、MRI和放射摄影多种成像方式。由放射科住院医师设计并标注问题,评估内容包括解剖结构识别、位置定位、侧别判断、区域划分、设备识别和结构间空间关系。评估流程支持标准化提示、结构化输出收集、响应归一化、大语言模型作为评判者打分、人工质量审核、二元正确性评分及按模态、解剖部位和推理类型进行子组分析。SPARC-Rad为评估视觉语言模型是否能将解剖结构视为一个空间系统提供了可复用框架,支持未来模型研发、故障模式分析与部署前评估。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, report generation, or broad visual question answering rather than the spatial and anatomical reasoning required for radiology. We developed the Spatial Perception and Anatomical Reasoning in Clinical Radiology (SPARC-Rad) Benchmark, a manually curated multimodal benchmark dataset and evaluation pipeline for assessing these capabilities in radiology VLMs. SPARC-Rad includes 300 image-question pairs derived from healthy control imaging studies in The Cancer Imaging Archive (TCIA), spanning CT, MRI, and radiography across the abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees manually designed and annotated questions to evaluate anatomical identification, localization, laterality, regional recognition, device identification, and inter-structure spatial relationships. The evaluation pipeline supports standardized prompting, structured output collection, response normalization, LLM-as-judge grading, human quality review, binary correctness scoring, and subgroup analysis by modality, anatomy, and reasoning type. SPARC-Rad provides a reusable framework for evaluating whether VLMs can provide reasoning for radiologic anatomy as a spatial system, supporting future model development, failure-mode analysis, and pre-deployment assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。