arXiv:2509.22258cs.CVcs.AI2025-09被引 4

提出神经科临床推理新基准,揭示现有模型在真实诊断中能力不足。

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks

  • 构建融合多序列MRI、病历和笔记的推理型评测集
  • 顶尖模型在新基准上表现大幅下降,错误主因是推理缺陷
  • 适合医疗AI可信性评估与临床级模型研发者使用

视觉语言模型在标准医学数据集上表现优异,但其真实临床推理能力仍不明确。现有数据集过度关注分类准确率,造成评估假象。本文提出Neural-MedBench,一个聚焦神经科临床推理的紧凑型评测基准,整合多序列MRI、结构化电子病历与临床文本,涵盖鉴别诊断、病灶识别与推理生成三类核心任务。采用结合大模型评分、医生验证与语义相似度的混合评分流程确保可靠性。对GPT-4o、Claude-4、MedGemma等先进模型的系统评估显示,其性能相比传统数据集显著下降,且错误以推理失败为主,非感知误差。研究强调需建立双轴评估框架:广度型大数据集用于统计泛化,深度型小规模基准如Neural-MedBench用于推理真实性检验。该基准已开源(https://neuromedbench.github.io/),支持未来评测扩展与低成本高可信度临床AI评估。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have achieved remarkable performance on standard medical benchmarks, yet their true clinical reasoning ability remains unclear. Existing datasets predominantly emphasize classification accuracy, creating an evaluation illusion in which models appear proficient while still failing at high-stakes diagnostic reasoning. We introduce Neural-MedBench, a compact yet reasoning-intensive benchmark specifically designed to probe the limits of multimodal clinical reasoning in neurology. Neural-MedBench integrates multi-sequence MRI scans, structured electronic health records, and clinical notes, and encompasses three core task families: differential diagnosis, lesion recognition, and rationale generation. To ensure reliable evaluation, we develop a hybrid scoring pipeline that combines LLM-based graders, clinician validation, and semantic similarity metrics. Through systematic evaluation of state-of-the-art VLMs, including GPT-4o, Claude-4, and MedGemma, we observe a sharp performance drop compared to conventional datasets. Error analysis shows that reasoning failures, rather than perceptual errors, dominate model shortcomings. Our findings highlight the necessity of a Two-Axis Evaluation Framework: breadth-oriented large datasets for statistical generalization, and depth-oriented, compact benchmarks such as Neural-MedBench for reasoning fidelity. We release Neural-MedBench at https://neuromedbench.github.io/ as an open and extensible diagnostic testbed, which guides the expansion of future benchmarks and enables rigorous yet cost-effective assessment of clinically trustworthy AI.

医疗AI临床推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。