arXiv:2603.25778cs.CV2026-03中稿 · CVPR

模仿医生看内镜视频的专注-感知过程,提升病灶识别能力

Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video Analysis

  • 分层建模:先聚焦病灶区域学静态特征,再感知其跨帧演化
  • 在11个数据集上超越现有方法,病灶检测准确率提升显著
  • 适合医疗视觉、医学影像分析研究者参考

内镜视频分析对胃肠道早期筛查至关重要,但受限于高质量标注数据稀缺。尽管自监督视频预训练展现潜力,现有方法多针对自然视频设计,过度强调密集时空建模与运动信息,忽视临床决策所需的静态结构语义。为此,我们提出聚焦-感知表征学习(FPRL),一种受认知启发的分层框架,模拟临床检查流程。FPRL首先聚焦于帧内病灶中心区域,学习静态语义;随后感知其跨帧演化,建模上下文语义。通过教师先验自适应掩码(TPAM)与多视图稀疏采样,捕捉静态语义,减少冗余时序依赖,强化病灶局部特征。接着利用跨视图掩码特征补全(CVMFC)与注意力引导时序预测(AGTP),建立跨视图对应关系,有效建模结构化帧间演化,增强时序语义连续性并保持全局上下文完整性。在11个内镜视频数据集上的广泛实验表明,FPRL在多种下游任务中表现优异,验证了其在内镜视频表征学习中的有效性。代码已开源。

原文摘要 · Abstract (English)

Endoscopic video analysis is essential for early gastrointestinal screening but remains hindered by limited high-quality annotations. While self-supervised video pre-training shows promise, existing methods developed for natural videos prioritize dense spatio-temporal modeling and exhibit motion bias, overlooking the static, structured semantics critical to clinical decision-making. To address this challenge, we propose Focus-to-Perceive Representation Learning (FPRL), a cognition-inspired hierarchical framework that emulates clinical examination. FPRL first focuses on intra-frame lesion-centric regions to learn static semantics, and then perceives their evolution across frames to model contextual semantics. To achieve this, FPRL employs a hierarchical semantic modeling mechanism that explicitly distinguishes and collaboratively learns both types of semantics. Specifically, it begins by capturing static semantics via teacher-prior adaptive masking (TPAM) combined with multi-view sparse sampling. This approach mitigates redundant temporal dependencies and enables the model to concentrate on lesion-related local semantics. Following this, contextual semantics are derived through cross-view masked feature completion (CVMFC) and attention-guided temporal prediction (AGTP). These processes establish cross-view correspondences and effectively model structured inter-frame evolution, thereby reinforcing temporal semantic continuity while preserving global contextual integrity. Extensive experiments on 11 endoscopic video datasets show that FPRL achieves superior performance across diverse downstream tasks, demonstrating its effectiveness in endoscopic video representation learning. The code is available at https://github.com/MLMIP/FPRL.

内镜分析自监督学习分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。