arXiv:2510.20967cs.CVcs.AI2025-10被引 5

首个支持3D医学影像逐步推理的标注数据集,助力AI理解临床诊断逻辑。

3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models

  • 构建3D膝关节MRI推理数据集,含49.4万条专家标注的五元组。
  • 每条数据包含定位框、诊断问题、推理步骤与结构化严重程度评估。
  • 为多模态医疗AI提供可解释的3D诊断决策测试基准,适合临床场景研究者。

当前视觉语言模型(VLMs)难以在3D医学图像中准确定位解剖区域并进行分步推理,而这正是真实诊断流程的核心需求。现有3D数据集虽提供定位标签,但缺乏对‘基于定位的推理’支持。为此,我们提出3DReasonKnee,首个面向医学影像的3D grounded reasoning数据集,涵盖7,970例3D膝关节MRI,生成494,000条高质量五元组。每个五元组包括:(1) 3D MRI体积,(2) 针对特定解剖区域的诊断问题,(3) 定位相关结构的3D边界框,(4) 临床医生生成的显式3D推理链,(5) 相关解剖区域的结构化严重程度评估。数据集创建与验证耗时超450小时专家工作,确保高质量与临床相关性。我们建立ReasonKnee-Bench以评估定位与诊断准确性,揭示VLM在不同解剖区域与诊断问题下的表现。我们对五种主流VLM进行基准测试,提供初始性能参考。3DReasonKnee作为骨科医生诊断经验的数字仓库,为推动多模态医疗AI实现3D、临床对齐、定位决策能力提供关键测试平台。数据集地址:https://huggingface.co/datasets/rajpurkarlab/3DReasonKnee

原文摘要 · Abstract (English)

Current Vision-Language Models (VLMs) struggle to ground anatomical regions in 3D medical images and reason about them in a step-by-step manner, a key requirement of real-world diagnostic assessment. This ability is essential for aligning model outputs with the diagnostic workflows clinicians use in practice, enabling trustworthy clinician-AI collaboration. Existing 3D datasets provide localization labels, but none support this "grounded reasoning" ability. To address this gap, we introduce 3DReasonKnee, the first 3D grounded reasoning dataset for medical images, which provides 494k high-quality quintuples derived from 7,970 3D knee MRI volumes. Each quintuple includes: (1) the 3D MRI volume, (2) a diagnostic question targeting a specific anatomical region (3) a 3D bounding box localizing the relevant anatomical structures, (4) clinician-generated diagnostic reasoning steps that explicitly detail the 3D reasoning process, and (5) structured severity assessments for the relevant anatomical region. The creation and validation of 3DReasonKnee, involving over 450 hours of expert clinician time for manually segmenting MRIs and generating reasoning chains, ensures its superior quality and clinical relevance. We establish ReasonKnee-Bench to evaluate localization and diagnostic accuracy, providing insight into VLM ability to perform grounding and severity assessment across anatomical regions and diagnostic inquiries. We benchmark five state-of-the-art VLMs, providing baseline performance for ReasonKnee-Bench. By providing this unique resource of expert-annotated 3D reasoning pathways, 3DReasonKnee serves as a repository of orthopedic surgeons' diagnostic expertise and offers a vital testbed for advancing multimodal medical AI systems towards 3D, clinically aligned, localized decision-making capabilities. The dataset can be found in: https://huggingface.co/datasets/rajpurkarlab/3DReasonKnee

医学影像3D推理多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。