arXiv:2511.01109cs.CV2025-11

将解剖先验融入视频变压器,提升超声心动图分析精度与可解释性。

Anatomically Constrained Transformers for Echocardiogram Analysis

  • 用点集表示心肌结构,结合空间几何与图像块编码为注意力令牌
  • 仅重建解剖区域图像块,使模型聚焦病灶区,提升对心衰的检测准确率
  • 无需额外组件即可实现心肌追踪,适合医疗影像分析研究者

视频变压器在超声心动图分析中展现出强大潜力,通过自监督预训练和灵活适配多种任务。然而,如同其他视频模型,它们容易从非诊断区域(如背景)学习虚假相关性。为此,我们提出视频解剖约束变压器(ViACT),将解剖先验直接嵌入变压器架构。ViACT将可变形解剖结构表示为点集,并将空间几何与对应图像块编码为变压器令牌。预训练阶段采用掩码自编码策略,仅对解剖区域图像块进行掩码与重建,强制表示学习聚焦于解剖区域。预训练模型可微调用于局部任务。本文聚焦心肌,应用于左室射血分数(EF)回归与心脏淀粉样变性(CA)检测任务。解剖约束使变压器注意力集中在心肌内,生成与已知CA病理区域对齐的可解释注意力图。此外,ViACT无需相关体积等专用组件,即可泛化至心肌点追踪任务。

原文摘要 · Abstract (English)

Video transformers have recently demonstrated strong potential for echocardiogram (echo) analysis, leveraging self-supervised pre-training and flexible adaptation across diverse tasks. However, like other models operating on videos, they are prone to learning spurious correlations from non-diagnostic regions such as image backgrounds. To overcome this limitation, we propose the Video Anatomically Constrained Transformer (ViACT), a novel framework that integrates anatomical priors directly into the transformer architecture. ViACT represents a deforming anatomical structure as a point set and encodes both its spatial geometry and corresponding image patches into transformer tokens. During pre-training, ViACT follows a masked autoencoding strategy that masks and reconstructs only anatomical patches, enforcing that representation learning is focused on the anatomical region. The pre-trained model can then be fine-tuned for tasks localized to this region. In this work we focus on the myocardium, demonstrating the framework on echo analysis tasks such as left ventricular ejection fraction (EF) regression and cardiac amyloidosis (CA) detection. The anatomical constraint focuses transformer attention within the myocardium, yielding interpretable attention maps aligned with regions of known CA pathology. Moreover, ViACT generalizes to myocardium point tracking without requiring task-specific components such as correlation volumes used in specialized tracking networks.

视频变压器超声心动图解剖先验可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。