用跨视图注意力建模心脏超声,提升多视角影像的表征能力。
Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography
- 在隐空间设计跨帧跨视图注意力,融合多视角心脏影像
- 在真实临床数据上预训练,实现对心脏病编码的视频预测
- 成人数据学的表征可有效迁移至儿童,泛化性强
超声心动图因其无创性和低成本被广泛用于心脏评估,但心脏的时空视图稀疏且异质,带来独特挑战。现有掩码自编码器通常独立处理图像或短片段,难以捕捉构建连贯心脏表征所需的多视图结构。本文提出针对医学影像多视图特性的基础模型——隐空间注意力掩码自编码器(LAMAE),在标准MAE基础上引入隐空间注意力模块,直接在潜在空间实现帧间与视图间的信息交换,可聚合不同长度序列与异构视图,从部分观测中重建心脏功能的整体表征。我们在大规模、未清理的真实世界数据集MIMIC-IV-ECHO上进行预训练。据我们所知,这是首个基于MIMIC-IV-ECHO视频预测ICD-10编码的结果。此外,实验表明,成人数据学习的表征能有效迁移至儿科群体,尽管解剖差异显著。结果证明,融入多视图结构先验可获得更鲁棒、更具迁移性的表示。
原文摘要 · Abstract (English)
Echocardiography is a widely used modality for cardiac assessment due to its non-invasive and cost-effective nature, but the sparse and heterogeneous spatiotemporal views of the heart pose distinct challenges. Existing masked autoencoder (MAE) approaches typically process images or short clips independently, failing to capture the inherent multi-view structure required for coherent cardiac representation. We introduce Latent Attention Masked Autoencoder (LAMAE), a foundation model architecture tailored to the multi-view nature of medical imaging. LAMAE augments the standard MAE with a latent attention module that enables information exchange across frames and views directly in latent space. This allows the model to aggregate variable-length sequences and distinct views, reconstructing a holistic representation of cardiac function from partial observations. We pretrain LAMAE on MIMIC-IV-ECHO, a large-scale, uncurated dataset reflecting real-world clinical variability. To the best of our knowledge, we present the first results for predicting ICD-10 codes from MIMIC-IV-ECHO videos. Furthermore, we empirically demonstrate that representations learned from adult data transfer effectively to pediatric cohorts despite substantial anatomical differences. These results provide evidence that incorporating structural priors, such as multi-view attention, yields significantly more robust and transferable representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。