用稀疏标注的运动信息融合多视角超声,提升心梗定位精度
Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography

- 仅需一个模板帧实现稀疏运动追踪,避免密集标注
- 融合运动与视觉特征,段级心梗定位F1达72.4%
- 适合超声影像分析、医学图像智能诊断的研究者
心肌梗死(MI)是全球主要致死原因。超声心动图(Echo)广泛用于评估,其中局部室壁运动异常是关键指标。现有方法多依赖手工特征或密集标注的运动估计,标注成本高。尽管基础模型提升了基于视觉的超声分析,但多数方法仅处理单视角,尤其在心尖视图下因视角依赖性导致定位不可靠。为此,我们提出MCF-Net,一种基于运动引导的多视角融合框架,将心肌运动线索与基础模型表征结合以定位梗死区域。使用预先训练的EchoPrime模型提取双视角视觉特征。心脏运动通过极稀疏监督建模:仅需一个标注模板帧即可跨视频初始化点追踪,避免密集标签。运动生成的段级软掩码提供粗略空间先验,选择性增强困难心肌段特征。运动条件融合机制在视角间整合运动与视觉信息,优化预测而不覆盖强外观线索。在段级心梗定位任务中,MCF-Net取得72.4% F1和84.9%准确率,优于当前最优的纯运动、纯视觉及融合基线。
原文摘要 · Abstract (English)
Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation, but the need for extensive annotation limits applicability. Foundation models have recently improved vision-based Echo analysis; however, most methods operate on single views and segment-level localization remains unreliable under view-dependent ambiguity, especially in apical views. To address this, we propose MCF-Net, a novel motion-guided multi-view fusion framework that fuses myocardial motion cues with foundation model representations to localize infarction. Visual features are extracted using EchoPrime, a pretrained Echo foundation model shared across dual views. Cardiac motion is modeled with extremely sparse supervision: a single annotated template frame is transferred across videos to initialize point tracking, avoiding dense labels. Motion-derived segment-aware soft masks provide coarse spatial priors that selectively enhance features for challenging myocardial segments. A motion-conditioned fusion mechanism then integrates motion and vision across views, refining predictions without overriding strong appearance cues. On segment-level MI localization, MCF-Net achieves 72.4\% F1 and 84.9\% accuracy, outperforming state-of-the-art motion-only, vision-only, and fusion baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。