arXiv:2607.07518cs.CV2026-07被引 1

融合形状与纹理特征,动态识别心脏关键阶段

Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification

论文配图:Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification
图 1 · 摘自论文原文
  • 在潜在空间设计双向交叉注意力,自适应融合形状与纹理特征
  • 在心腔磁共振视频数据集上达到顶尖分类性能
  • 可解释性强,精准定位诊断关键时相与模态贡献

可变形形状表示在心脏图像分类中已被证明是纹理特征的稳健补充,提供对成像伪影和强度变化不变的几何先验。然而,现有深度网络仅通过简单拼接方式融合这些不同特征表示,未能充分挖掘其互补性,也未学习跨模态特征依赖关系。此外,该方法对所有时间点施加统一注意力,忽略了心脏各相位间诊断重要性的差异。本文提出一种新型心脏视频分类模型,首次在可变形形状与图像纹理表示的联合空间中学习时序特征。我们设计了潜空间中的双向交叉注意力机制,融合潜在可变形形状与图像特征,使每种模态能基于时空对应关系自适应地加权另一模态。相比当前方法对所有心脏相位采用统一权重,本方法学习动态调整来自图像的形状与纹理表示贡献。我们在一个心动电影心脏磁共振(CMR)视频数据集上验证了该方法的先进分类性能,并通过注意力机制提升了可解释性,准确识别出具有诊断意义的心脏关键相位及各模态贡献。

原文摘要 · Abstract (English)

Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors that are invariant to imaging artifacts and intensity variations. However, existing deep networks perform simple concatenation to combine these distinct feature representations, which neither fully exploits their complementary nature nor learns cross-modal feature dependencies. Furthermore, this results in uniform attention across all timepoints; hence ignoring the varying diagnostic importance across the cardiac phases. In this paper, we propose a novel cardiac video classification model that, for the first time, learns temporal features in an integrated space of deformable shape and image texture representations. In particular, we design a bi-directional cross-attention in the latent space to fuse latent deformable shape and image features, allowing each modality to adaptively weight the other based on spatio-temporal correspondence. In contrast to current methods that apply uniform weighting across all the cardiac phases, our approach learns to dynamically adjust the contributions of shape and texture representations, derived from images, over time. We demonstrate state-of-the-art classification performance on a cine cardiac magnetic resonance (CMR) video dataset, achieving improved interpretability from attention mechanisms that identify diagnostically critical cardiac phases and modality contributions.

心脏影像多模态融合动态注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。