构建脑肿瘤MRI多模态诊断数据集,推动模型生成临床可解释推理。
MM-NeuroOnco: A Multimodal Benchmark and Instruction Dataset for MRI-Based Brain Tumor Diagnosis
- 设计多模型协作管道,自动生成丰富语义标注,突破传统掩码标注局限。
- 涵盖24,726张MRI切片与20万条指令,评估显示最强基线模型准确率仅41.88%。
- 适合医学影像、多模态推理与AI辅助诊断研究者使用。
精准脑肿瘤诊断需模型不仅检测病灶,还需基于影像表现生成临床可解释的推理,但现有公开数据集在标注丰富度和诊断语义上仍显不足。为此,我们提出MM-NeuroOnco,一个大规模多模态基准与指令微调数据集,包含来自20个数据源的24,726张脑肿瘤MRI切片,配以约20万条语义丰富的多模态指令,覆盖多种肿瘤亚型与成像模态。为缓解诊断语义标注稀缺与高成本问题,我们开发了多模型协同的自动化医疗信息补全与质量控制流程,实现超越仅掩码标注的诊断语义生成。基于该数据集,我们进一步构建了人工标注的评估基准MM-NeuroOnco-Bench,采用拒绝感知设置以减少封闭式问答带来的偏差。对十种代表性模型的评估显示,即使最强基线Gemini 3 Flash在诊断相关问题上也仅达41.88%准确率,凸显多模态脑肿瘤诊断理解的挑战性。利用该数据集,我们提出NeuroOnco-GPT,经微调后诊断问题准确率提升27个百分点。结果表明,本数据集与基准能有效推动临床导向的多模态诊断推理发展。代码与数据集已公开于:https://github.com/gfnnnb/MM-NeuroOnco
原文摘要 · Abstract (English)
Accurate brain tumor diagnosis requires models to not only detect lesions but also generate clinically interpretable reasoning grounded in imaging manifestations, yet existing public datasets remain limited in annotation richness and diagnostic semantics. To bridge this gap, we introduce MM-NeuroOnco, a large-scale multimodal benchmark and instruction-tuning dataset for brain tumor MRI understanding, consisting of 24,726 MRI slices from 20 data sources paired with approximately 200,000 semantically enriched multimodal instructions spanning diverse tumor subtypes and imaging modalities. To mitigate the scarcity and high cost of diagnostic semantic annotations, we develop a multi-model collaborative pipeline for automated medical information completion and quality control, enabling the generation of diagnosis-related semantics beyond mask-only annotations. Building upon this dataset, we further construct MM-NeuroOnco-Bench, a manually annotated evaluation benchmark with a rejection-aware setting to reduce biases inherent in closed-ended question formats. Evaluation across ten representative models shows that even the strongest baseline, Gemini 3 Flash, achieves only 41.88% accuracy on diagnosis-related questions, highlighting the substantial challenges of multimodal brain tumor diagnostic understanding. Leveraging MM-NeuroOnco, we further propose NeuroOnco-GPT, which achieves a 27% absolute accuracy improvement on diagnostic questions following fine-tuning. This result demonstrates the effectiveness of our dataset and benchmark in advancing clinically grounded multimodal diagnostic reasoning. Code and dataset are publicly available at: https://github.com/gfnnnb/MM-NeuroOnco
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。