arXiv:2603.17879cs.CVcs.AI2026-03

针对胃肠道内镜视频罕见病灶分类,提出解剖引导的视觉语言模型与角度分离损失。

Anatomy-Guided Vision-Language Learning with Angular Prototype Separation for Multi-Label Video Capsule Endoscopy Classification Under Class Imbalance

  • 用角度分离损失防止稀有类别原型坍塌,提升小样本病灶识别能力。
  • 引入生物状态机解码器,将每视频输出从数百个伪事件压缩至2-3个真实事件。
  • 在16万帧测试集上,检测准确率较旧方法提升44%-46%,单卡仅需21分钟推理。

本文提出一种用于视频胶囊内镜(VCE)的多标签时序事件检测框架,针对Galar数据集中的极端类别不平衡问题,结合两大核心贡献:类原型的角度分离损失与基于生理学的生物状态机时序解码器。采用BiomedCLIP作为骨干模型,通过局部差分注意力模块融合三帧图像,增强瞬态病灶信号并抑制静态冗余。额外设计解剖上下文头,利用胃肠道病变的空间共现结构对病灶预测进行软解剖条件约束。可学习的文本提示与原型级逻辑增强,配合惩罚类原型间非对角余弦相似度的角度分离损失,有效缓解稀有类别在极端不平衡下的原型坍塌。训练策略融合非对称焦点损失、逆频加权采样、时序Mixup、指数移动平均与每类阈值校准。生物状态机解码器取代简单的间隙合并,采用基于解剖标签的前向唯一状态转移机制,消除此前导致每视频产生数百个虚假事件的碎片化伪影,使每视频解剖事件数降至2–3个临床合理范围。在包含三个NaviCam检查的独立测试集RARE-VISION(共161,025帧)上,新流程实现[email protected]为0.3597,[email protected]为0.3399,相较先前提交结果分别提升46%和44%,单卡推理耗时约21分钟。

原文摘要 · Abstract (English)

This work presents a multi-label temporal event detection framework for video capsule endoscopy (VCE) that addresses the extreme class imbalance inherent in the Galar dataset by combining two principal contributions: an Angular Separation Loss on class prototypes and a Biological State Machine temporal decoder. The backbone remains BiomedCLIP, a biomedical vision-language foundation model. Three consecutive frames are fused through a Local Differencing Attention module that amplifies transient pathological signals by suppressing static temporal redundancy. An Anatomy Context Head then conditions pathological predictions on soft anatomical activations, exploiting the known spatial co-occurrence structure of GI findings. Learnable text-feature prompts and prototype-based logit augmentation are trained alongside an Angular Separation Loss that penalizes off-diagonal cosine similarity between class prototypes, preventing the prototype collapse that afflicts rare classes under extreme imbalance. To counteract the skewed label distribution, the training regime combines asymmetric focal loss, inverse-frequency weighted sampling, temporal Mixup, Exponential Moving Average, and per-class threshold calibration. The Biological State Machine decoder replaces naive gap merging with a physiologically grounded forward-only state transition over anatomy labels, eliminating the fragmentation artefact that produced hundreds of spurious anatomy events per video in the prior approach and reducing per-video anatomy output to 2--3 clinically realistic events. On the held-out RARE-VISION test set comprising three NaviCam examinations (161,025 frames), the updated pipeline achieves an overall temporal [email protected] of 0.3597 and [email protected] of 0.3399, representing a relative improvement of 46% and 44% respectively over the prior submission, with total inference completed in approximately 21 minutes on a single GPU.

视频分类医学影像多标签检测小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。