统一建模多视角超声心动图关键帧,提升检测准确率与跨视角泛化能力
FrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection

- 分层运动建模:视内多任务学习减少外观偏差,视间共享运动特征
- 在25,872段视频上实现最优关键帧检测精度,跨视角泛化性能强
- 适用于多视图超声心动图分析,尤其适合缺乏标注数据的场景
准确检测收缩末期(ES)和舒张末期(ED)帧是超声心动图评估的基础。现有方法通常针对特定视角设计,依赖辅助标注或复杂的视觉建模,限制了泛化能力。在多视角建模中,关键帧检测依赖于共享的心脏运动规律,但显著的外观差异和运动模式变化使得统一建模极具挑战。为此,我们提出FrameONE,一个面向多视角超声心动图关键帧检测的统一端到端框架。FrameONE引入分层运动建模策略:视内多任务学习降低外观偏差,促进各视角内的运动聚焦表示;视间通用运动学习模块进一步分离视图无关动态与视图特异性模式,实现跨视图的共享且灵活的运动表示学习。在涵盖四个标准视图的25,872段视频上进行的大量实验表明,FrameONE在关键帧检测精度上达到当前最优水平,并展现出强大的跨视角泛化能力。代码已公开于https://github.com/szuboy/FrameONE。
原文摘要 · Abstract (English)
Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically developed in a view-specific manner, depend on auxiliary annotations or intensive visual modeling, which limits their generalizability. In multi-view modeling, keyframe detection is driven by shared cardiac motion, yet large appearance differences and motion patterns make unified modeling challenging. To address these issues, we propose FrameONE, a unified end-to-end framework for multi-view echocardiographic keyframe detection. FrameONE introduces a Hierarchical Motion Modeling strategy: an intra-view multi-task learning reduces appearance bias and promotes motion-focused representations within each view; an inter-view general motion learning module further separates view-agnostic dynamics from view-specific patterns, enabling shared yet flexible motion representation learning across views. Extensive experiments on 25,872 videos spanning four standard views demonstrate that FrameONE achieves state-of-the-art keyframe detection accuracy with strong cross-view generalization. Code is available at https://github.com/szuboy/FrameONE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。