arXiv:2603.02790cs.CV2026-03被引 4

构建统一医学基准UNICORN,评估跨模态医疗模型性能。

Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language

  • 采用两步框架分离模型推理与任务评估,确保公平对比。
  • 涵盖2400+患者、3700+影像和2400+病历报告,覆盖8个解剖区、4种模态。
  • 提供标准化评分与开源平台,适合评估医疗大模型泛化能力。

医疗基础模型有望从大规模多样化数据中学习通用特征,实现跨模态泛化与少量样本快速适应新任务。然而,现有公开评估框架碎片化,缺乏统一标准,限制了跨任务泛化能力的验证。本文提出UNICORN,一个面向计算病理学、放射学与自然语言的统一基准,采用新型两步框架,将模型推理与任务特定评估解耦,通过间接访问的隔离测试集(来自17所机构、8个国家的临床队列)与标准化评估代码,支持可复现评估。测试集包含超过2,400名患者的多源数据,涵盖3,700余例影像与2,400余份临床报告,覆盖8个解剖区域与4种成像模态。引入统一的UNICORN Score,支持跨领域、跨模态、跨任务的直接比较。提供任务级与聚合排行榜,所有数据、基线方法与评估平台均开放于unicorn.grand-challenge.org。

原文摘要 · Abstract (English)

Medical foundation models show promise to learn broadly generalizable features from large, diverse datasets. This could be the base for reliable cross-modality generalization and rapid adaptation to new, task-specific goals, with only a few task-specific examples. Yet, evidence for this is limited by the lack of public, standardized, and reproducible evaluation frameworks, as existing public benchmarks are often fragmented across task-, organ-, or modality-specific settings, limiting assessment of cross-task generalization. We introduce UNICORN, a public benchmark designed to systematically evaluate medical foundation models under a unified protocol. To isolate representation quality, we built the benchmark on a novel two-step framework that decouples model inference from task-specific evaluation based on standardized few-shot adaptation. As a central design choice, we constructed indirectly accessible sequestered test sets derived from clinically relevant cohorts, along with standardized evaluation code and a submission interface on an open benchmarking platform. Performance is aggregated into a single UNICORN Score, a new metric that we introduce to support direct comparison of foundation models across diverse medical domains, modalities, and task types. The UNICORN test dataset includes data from more than 2,400 patients, including over 3,700 vision cases and over 2,400 clinical reports collected from 17 institutions across eight countries. The benchmark spans eight anatomical regions and four imaging modalities. Both task-specific and aggregated leaderboards enable accessible, standardized, and reproducible evaluation. By standardizing multi-task, multi-modality assessment, UNICORN establishes a foundation for reproducible benchmarking of medical foundation models. Data, baseline methods, and the evaluation platform are publicly available via unicorn.grand-challenge.org.

医疗大模型多模态基准评测可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。