融合时空模型与解剖感知,提升罕见病内镜事件检测精度
VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

- 采用双骨干网络捕捉时间上下文与帧级视觉语义
- 后处理阶段通过验证引导加权融合与解剖感知解码,提升事件检测性能
- 在罕见病内镜检测任务中达到0.3726的[email protected],适合医学影像分析研究者参考
胶囊内镜事件检测因临床相关发现稀疏、视觉异质性强,且评估以事件为单位而非帧级准确率而极具挑战。本文提出VISTA,一种面向RAREVISION任务的度量对齐多骨干框架。VISTA结合EndoFM-LV获取时间上下文,DINOv3 ViTL/16提取帧级视觉语义,随后通过多样头集成(DHE)、验证引导加权融合(VGWF)和解剖感知时间事件解码(ATED)进行联合优化。初始提交在隐藏测试集上取得0.3530的[email protected]和0.3235的[email protected];赛后通过局部阈值优化与全局粗搜索扩展,性能提升至0.3726 [email protected]和0.3431 [email protected],团队在赛后评估中排名第二。
原文摘要 · Abstract (English)
Capsule endoscopy event detection is challenging because clinically relevant findings are sparse, visually heterogeneous, and evaluated at the event level rather than by frame accuracy. We propose VISTA, a metric-aligned multi-backbone framework for the RAREVISION task. VISTA combines EndoFM-LV for temporal context and DINOv3 ViTL/16 for frame-level visual semantics, followed by a Diverse Head Ensemble (DHE), Validation-Guided Weighted Fusion (VGWF), and Anatomy-Aware Temporal Event Decoding (ATED). The original official submission achieved hidden-test temporal [email protected] of 0.3530 and [email protected] of 0.3235. After the competition, extending local threshold refinement with a global coarse search improved performance to 0.3726 [email protected] and 0.3431 [email protected], ranking Team ACVLab second in the post-competition evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。