针对胶囊内镜视频的复杂分类任务,提出双分支时序模型提升多病灶定位精度。
GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

- 分层建模结合双图卷积与双向Mamba,捕捉局部细节与长程依赖
- 改进后在RARE-VISION数据集上[email protected]达0.3409,[email protected]达0.3333
- 特别适合处理器官与病灶混杂、罕见类别难识别的医学视频分析
视频胶囊内镜(VCE)面临一个复杂的多标签时序分类问题,需在数万帧中同时定位8个解剖区域并检测9种病理发现。我们提出GALAR-TemporalNet v2,一种分层时序模型,解决三个核心挑战:极端类别不平衡、长程时序依赖以及病理与解剖信息纠缠。该架构结合窗口自注意力进行局部建模,双图卷积网络(Dual-Graph GCN)捕捉全局帧间关系,以及双向Mamba实现选择性边界上下文编码。新颖的解剖原型残差路径将病理异常信号与正常器官外观分离,帧级GCN捷径连接则稳定了视觉相似的稀有类别的训练。竞赛版本GALAR-TemporalNet在RARE-VISION测试集上取得[email protected]为0.2644、[email protected]为0.2353的成绩。赛后重构的GALAR-TemporalNet v2通过优化病理分支结构、改进损失函数并扩展后处理,将结果提升至[email protected] 0.3409,[email protected] 0.3333。
原文摘要 · Abstract (English)
Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regions and detection of 9 pathological findings across tens of thousands of frames. We present GALAR-TemporalNet v2, a hierarchical temporal model that addresses three core challenges: extreme class imbalance, long-range temporal dependencies, and pathology--anatomy entanglement. Our architecture combines windowed self-attention for local modeling, a Dual-Graph GCN for global frame relationships, and Bidirectional Mamba for selective boundary context encoding. A novel anatomy prototype residual pathway decouples pathological deviation signals from normal organ appearance, and a frame-level GCN skip connection stabilizes training of visually confusable rare classes. The competition version, GALAR-TemporalNet, achieved an overall [email protected] of 0.2644 and [email protected] of 0.2353 on the RARE-VISION test set. Following the competition, the redesigned GALAR-TemporalNet v2 -- incorporating a restructured pathology branch, refined loss functions, and extended post-processing -- improved these results to [email protected] of 0.3409 and [email protected] of 0.3333.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。