首个标准化超声心动图大模型评测基准,验证通用模型泛化能力
CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?
- 构建统一数据集框架,覆盖4类回归与5类分类任务
- 通用编码器在零样本下表现良好,但细粒度识别能力弱
- 捕捉心脏动态的模型更适合功能评估,检索方法泛化更稳
基础模型正重塑医学影像领域,但在超声心动图中的应用受限于对私有数据集的依赖,导致难以复现对比。超声心动图存在采集噪声大、帧间冗余高、公开数据集少等挑战。为此,我们提出CardioBench,一个面向超声心动图基础模型的综合性评测基准。该基准将8个公开数据集整合为标准化套件,涵盖4项回归和5项分类任务,覆盖功能、结构、诊断及视图识别等目标。基于此框架,我们在一致的零样本、探针和对齐协议下评估多个领先基础模型,包括心脏专用、生物医学及通用编码器。分析显示,通用编码器虽能较好迁移且接近探针性能,但在视图分类和细微病灶识别上表现显著不足;能捕捉心脏时序动态的模型在功能任务上最优,而基于检索的方法跨数据集泛化更稳定。通过发布预处理流程、数据划分和公开评估管道,CardioBench建立了可复现的参考标准,助力未来超声心动图及其他医学影像基础模型的架构设计。
原文摘要 · Abstract (English)
Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private datasets that prevent reproducible comparison. Echocardiography poses unique challenges, including noisy acquisitions, high frame redundancy, and limited diverse public datasets. To address this, we introduce CardioBench, a comprehensive benchmark for echocardiography foundation models. Specifically, CardioBench unifies eight publicly available datasets into a standardized suite spanning four regression and five classification tasks, covering functional, structural, diagnostic, and view recognition endpoints. Leveraging this framework, we evaluate several leading foundation models, including cardiac-specific, biomedical, and general-purpose encoders, under consistent zero-shot, probing, and alignment protocols. Our analysis reveals that while general-purpose encoders transfer well and often close the gap with probing, they struggle significantly with fine-grained distinctions like view classification and subtle pathology recognition. Results indicate that models capturing temporal cardiac dynamics perform best on functional tasks, while retrieval-based approaches generalize more consistently across datasets. By releasing preprocessing, splits, and public evaluation pipelines, CardioBench establishes a reproducible reference point to guide the architectural design of future echocardiography and possibly other medical imaging foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。