对比三种模型在心脏超声图像上的分割性能,给出可复现的基准测试和数据处理建议。
Unified Review and Benchmark of Deep Segmentation Architectures for Cardiac Ultrasound on CAMUS
- 统一标准下比较U-Net、Attention U-Net与TransUNet在CAMUS数据集上的表现。
- 原生NIfTI数据训练的U-Net达94%平均Dice,PNG-16bit为91%。
- 自监督预训练与GPT伪标签可提升泛化性,适合医疗影像研究者参考。
本文结合对心脏超声分割文献的聚焦综述与在CAMUS超声数据集上对三种代表性架构——U-Net、Attention U-Net与TransUNet的受控对比。实验涵盖多种预处理路径:原始NIfTI体积、16位PNG导出、GPT辅助的多边形伪标签,以及数千帧无标签动态序列的自监督预训练(SSL)。在相同训练划分、损失函数与评估标准下,直接在原始NIfTI数据上训练的U-Net取得94%的平均Dice系数;而16位PNG流程则达到91%。Attention U-Net在小目标或低对比区域表现略优,减少边界泄漏;TransUNet因能建模全局空间上下文,在复杂帧上展现最强泛化能力,尤其在初始化使用SSL时。伪标签经置信度过滤后扩大了训练集并提升了鲁棒性。本研究贡献包括:在标准化预处理与评估下对三类模型的可复现基准;数据准备中保持强度保真、分辨率一致性和配准对齐的实用指导;以及关于可扩展自监督与新兴多模态GPT驱动标注流水线的展望,适用于快速标注、质量控制与针对性数据集构建。
原文摘要 · Abstract (English)
Several review papers summarize cardiac imaging and DL advances, few works connect this overview to a unified and reproducible experimental benchmark. In this study, we combine a focused review of cardiac ultrasound segmentation literature with a controlled comparison of three influential architectures, U-Net, Attention U-Net, and TransUNet, on the Cardiac Acquisitions for Multi-Structure Ultrasound Segmentation (CAMUS) echocardiography dataset. Our benchmark spans multiple preprocessing routes, including native NIfTI volumes, 16-bit PNG exports, GPT-assisted polygon-based pseudo-labels, and self-supervised pretraining (SSL) on thousands of unlabeled cine frames. Using identical training splits, losses, and evaluation criteria, a plain U-Net achieved a 94% mean Dice when trained directly on NIfTI data (preserving native dynamic range), while the PNG-16-bit workflow reached 91% under similar conditions. Attention U-Net provided modest improvements on small or low-contrast regions, reducing boundary leakage, whereas TransUNet demonstrated the strongest generalization on challenging frames due to its ability to model global spatial context, particularly when initialized with SSL. Pseudo-labeling expanded the training set and improved robustness after confidence filtering. Overall, our contributions are threefold: a harmonized, apples-to-apples benchmark of U-Net, Attention U-Net, and TransUNet under standardized CAMUS preprocessing and evaluation; practical guidance on maintaining intensity fidelity, resolution consistency, and alignment when preparing ultrasound data; and an outlook on scalable self-supervision and emerging multimodal GPT-based annotation pipelines for rapid labeling, quality assurance, and targeted dataset curation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。