arXiv:2512.12492cs.CVcs.CL2025-12被引 4

针对内镜图像中假阴性高的问题,提出自适应检测-验证框架,显著提升零样本场景下息肉检出率。

Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings

  • 采用双阶段架构:检测器自适应调整置信度阈值,验证器通过强化学习优化
  • 在合成退化数据上测试,召回率提升14至22个百分点,误检率控制在基准附近
  • 特别适合临床场景,减少漏诊风险,对内镜医生辅助诊断有实际价值

在真实内镜场景中,光照变化、运动模糊和遮挡会严重降低图像质量,导致在清洁数据集上训练的息肉检测器性能下降。现有方法难以应对实验室与临床实践之间的领域差距。本文提出AdaptiveDetector,一种基于YOLOv11检测器与视觉语言模型(VLM)验证器的两阶段框架。检测器在VLM引导下自适应调整每帧置信度阈值,验证器则通过组相对策略优化(GRPO)和非对称成本敏感奖励函数进行微调,以抑制漏检——这是临床关键需求。为实现真实评估,我们系统性地对清洁数据集施加临床常见退化条件,构建综合性合成测试平台,用于零样本评测。在合成退化的CVC-ClinicDB和Kvasir-SEG图像上进行大量零样本测试表明,本方法相比YOLO单独使用,召回率提升14至22个百分点,精度波动在基准值上下0.7至1.7个百分点之间。自适应阈值与成本敏感强化学习相结合,实现了临床对齐的开放世界息肉检测,显著减少假阴性,降低癌前病变漏诊风险,改善患者预后。

原文摘要 · Abstract (English)

Polyp detectors trained on clean datasets often underperform in real-world endoscopy, where illumination changes, motion blur, and occlusions degrade image quality. Existing approaches struggle with the domain gap between controlled laboratory conditions and clinical practice, where adverse imaging conditions are prevalent. In this work, we propose AdaptiveDetector, a novel two-stage detector-verifier framework comprising a YOLOv11 detector with a vision-language model (VLM) verifier. The detector adaptively adjusts per-frame confidence thresholds under VLM guidance, while the verifier is fine-tuned with Group Relative Policy Optimization (GRPO) using an asymmetric, cost-sensitive reward function specifically designed to discourage missed detections -- a critical clinical requirement. To enable realistic assessment under challenging conditions, we construct a comprehensive synthetic testbed by systematically degrading clean datasets with adverse conditions commonly encountered in clinical practice, providing a rigorous benchmark for zero-shot evaluation. Extensive zero-shot evaluation on synthetically degraded CVC-ClinicDB and Kvasir-SEG images demonstrates that our approach improves recall by 14 to 22 percentage points over YOLO alone, while precision remains within 0.7 points below to 1.7 points above the baseline. This combination of adaptive thresholding and cost-sensitive reinforcement learning achieves clinically aligned, open-world polyp detection with substantially fewer false negatives, thereby reducing the risk of missed precancerous polyps and improving patient outcomes.

医学图像零样本检测强化学习内镜分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。