EvalBlocks让医学影像大模型评估更高效,一键完成多模型对比。
EvalBlocks: A Modular Pipeline for Rapidly Evaluating Foundation Models in Medical Imaging
- 模块化设计,可快速接入新数据集与模型
- 支持5个前沿模型在3项任务上快速评估
- 适合需要频繁测试模型的医学影像研究者
医学影像领域的大模型开发需持续监控下游性能。研究人员常面临实验繁杂、设计选择众多且依赖手动、非标准化流程的问题,效率低且易出错。本文提出 EvalBlocks,一个基于 Snakemake 构建的模块化、即插即用框架,用于高效评估大模型开发过程中的表现。该框架可无缝集成新数据集、大模型、聚合方法和评估策略,所有实验与结果集中追踪,单命令即可复现。通过高效缓存与并行执行,可在共享计算资源上规模化使用。在5个前沿大模型和3个医学影像分类任务上验证,显著简化评估流程,使研究者能更快迭代,专注模型创新。框架已开源,地址:https://github.com/DIAGNijmegen/eval-blocks。
原文摘要 · Abstract (English)
Developing foundation models in medical imaging requires continuous monitoring of downstream performance. Researchers are burdened with tracking numerous experiments, design choices, and their effects on performance, often relying on ad-hoc, manual workflows that are inherently slow and error-prone. We introduce EvalBlocks, a modular, plug-and-play framework for efficient evaluation of foundation models during development. Built on Snakemake, EvalBlocks supports seamless integration of new datasets, foundation models, aggregation methods, and evaluation strategies. All experiments and results are tracked centrally and are reproducible with a single command, while efficient caching and parallel execution enable scalable use on shared compute infrastructure. Demonstrated on five state-of-the-art foundation models and three medical imaging classification tasks, EvalBlocks streamlines model evaluation, enabling researchers to iterate faster and focus on model innovation rather than evaluation logistics. The framework is released as open source software at https://github.com/DIAGNijmegen/eval-blocks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。