arXiv:2605.22096cs.CV2026-05

融合时空模型与解剖感知,提升罕见病内镜事件检测精度

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

论文配图:VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results
图 1 · 摘自论文原文
  • 采用双骨干网络捕捉时间上下文与帧级视觉语义
  • 后处理阶段通过验证引导加权融合与解剖感知解码,提升事件检测性能
  • 在罕见病内镜检测任务中达到0.3726的[email protected],适合医学影像分析研究者参考

胶囊内镜事件检测因临床相关发现稀疏、视觉异质性强,且评估以事件为单位而非帧级准确率而极具挑战。本文提出VISTA,一种面向RAREVISION任务的度量对齐多骨干框架。VISTA结合EndoFM-LV获取时间上下文,DINOv3 ViTL/16提取帧级视觉语义,随后通过多样头集成(DHE)、验证引导加权融合(VGWF)和解剖感知时间事件解码(ATED)进行联合优化。初始提交在隐藏测试集上取得0.3530的[email protected]和0.3235的[email protected];赛后通过局部阈值优化与全局粗搜索扩展,性能提升至0.3726 [email protected]和0.3431 [email protected],团队在赛后评估中排名第二。

原文摘要 · Abstract (English)

Capsule endoscopy event detection is challenging because clinically relevant findings are sparse, visually heterogeneous, and evaluated at the event level rather than by frame accuracy. We propose VISTA, a metric-aligned multi-backbone framework for the RAREVISION task. VISTA combines EndoFM-LV for temporal context and DINOv3 ViTL/16 for frame-level visual semantics, followed by a Diverse Head Ensemble (DHE), Validation-Guided Weighted Fusion (VGWF), and Anatomy-Aware Temporal Event Decoding (ATED). The original official submission achieved hidden-test temporal [email protected] of 0.3530 and [email protected] of 0.3235. After the competition, extending local threshold refinement with a global coarse search improved performance to 0.3726 [email protected] and 0.3431 [email protected], ranking Team ACVLab second in the post-competition evaluation.

医学图像事件检测时空建模罕见病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。