arXiv:2603.18343cs.CV2026-03

用解剖约束和验证引导融合,提升罕见病胶囊内镜事件检测精度

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection

  • 融合局部时序与全局视觉模型,通过验证集优化权重
  • 在隐藏测试集上达到0.3530的[email protected]和0.3235的[email protected]
  • 适合医疗视频事件检测、罕见病智能诊断研究者参考

胶囊内镜事件检测因病灶稀疏、视觉异质且嵌入长而嘈杂的视频流中而极具挑战性,且评估以事件级为准而非帧级准确率。因此,我们将该任务定义为与指标对齐的事件检测问题,而非纯帧分类。框架结合两个互补主干:用于局部时序上下文的EndoFM-LV和用于强帧级视觉语义的DINOv3 ViT-L/16,随后通过多样头集成、验证引导的分层融合及解剖感知的时间事件解码实现整合。融合阶段采用验证集导出的类别级模型加权、主干加权和概率校准;解码阶段应用时间平滑、解剖约束、阈值优化和每标签事件生成,以获得稳定事件预测。消融实验表明,互补主干、验证引导融合与解剖感知解码均显著提升事件级性能。在官方隐藏测试集上,方法整体达到0.3530的[email protected]和0.3235的[email protected]

原文摘要 · Abstract (English)

Capsule endoscopy event detection is challenging because diagnostically relevant findings are sparse, visually heterogeneous, and embedded in long, noisy video streams, while evaluation is performed at the event level rather than by frame accuracy alone. We therefore formulate the RARE-VISION task as a metric-aligned event detection problem instead of a purely frame-wise classification task. Our framework combines two complementary backbones, EndoFM-LV for local temporal context and DINOv3 ViT-L/16 for strong frame-level visual semantics, followed by a Diverse Head Ensemble, Validation-Guided Hierarchical Fusion, and Anatomy-Aware Temporal Event Decoding. The fusion stage uses validation-derived class-wise model weighting, backbone weighting, and probability calibration, while the decoding stage applies temporal smoothing, anatomical constraints, threshold refinement, and per-label event generation to produce stable event predictions. Validation ablations indicate that complementary backbones, validation-guided fusion, and anatomy-aware temporal decoding all contribute to event-level performance. On the official hidden test set, the proposed method achieved an overall temporal [email protected] of 0.3530 and temporal [email protected] of 0.3235.

医疗影像事件检测多模态融合罕见病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。