arXiv:2606.22943cs.CV2026-06

评估自监督医学影像表示时,下游策略影响结果,需多方法验证。

Evaluating self-supervised echocardiographic representations across downstream extraction strategies for left-ventricular segmentation and ejection fraction estimation

论文配图:Evaluating self-supervised echocardiographic representations across downstream extraction strategies for left-ventricular segmentation and ejection fraction estimation
图 1 · 摘自论文原文
  • 用多种提取策略系统测试自监督特征的潜力
  • 轻量解码器比传统方法提升分割与射血分数预测性能
  • 结论依赖下游方法,单一评估易误导

自监督学习(SSL)在医学影像中可减少标注需求,但表征质量常依赖单一下游任务评估。对于密集临床任务,这可能将表征质量与下游模型能力混淆。本文在EchoNet-Dynamic数据集上系统评估了左心室分割与射血分数(EF)估计中两类自监督表征:通用冻结的DINOv3和任务适配的BYOS。比较了从启发式提取到部分微调的多层级提取策略。结果显示,仅用启发式提取时,DINOv3的Dice为0.684,EF MAE为13.01;而使用冻结轻量解码器后,性能提升至Dice 0.906,EF MAE 9.65,接近监督U-Net基线(Dice 0.915,EF MAE 9.72)。BYOS同样从Dice 0.687、EF MAE 17.83提升至Dice 0.902、EF MAE 8.74。表明自监督表征质量评估高度依赖下游策略,建议采用多策略评估。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) is increasingly used in medical imaging to reduce annotation requirements, but representation quality is often judged using a single downstream evaluation setting. For dense clinical tasks, this can confound representation quality with the capacity of the downstream model used to recover task-relevant information. We present a systematic evaluation of self-supervised representations for left-ventricular segmentation and ejection fraction (EF) estimation from apical four-chamber echocardiography on EchoNet-Dynamic. Rather than relying on a single downstream probe, we compare a hierarchy of extraction strategies with increasing expressivity: heuristic extraction without mask-supervised training, frozen linear probes, frozen lightweight decoder probes, and partial fine-tuning. We apply this framework to two complementary representation families: generic frozen self-DIstillation with NO labels (DINOv3) features and a task-adapted dense self-supervised representation, Bootstrap Your Own Segmentation (BYOS). In both families, heuristic extraction substantially understated what was recoverable from the frozen representation. For DINOv3, performance improved from Dice 0.684 and EF mean absolute error (MAE) 13.01 under heuristic extraction to Dice 0.906 and EF MAE 9.65 with a frozen lightweight decoder, approaching a supervised U-Net baseline (Dice 0.915, EF MAE 9.72). For BYOS, performance improved from Dice 0.687 and EF MAE 17.83 under heuristic extraction to Dice 0.902 and EF MAE 8.74 with a frozen lightweight decoder. These results show that conclusions about self-supervised representation quality in dense echocardiographic analysis depend strongly on the downstream extraction strategy used for evaluation. We therefore argue that multi-strategy evaluation is an important methodological consideration for SSL in dense medical image analysis.

自监督学习超声心动图左心室分割表征评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。