arXiv:2606.30951cs.CVcs.AI2026-06中稿 · MICCAI 2026

用强化学习教会AI在超声图像中找关键位置,提升前列腺癌检测准确率。

Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection

论文配图:Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection
图 1 · 摘自论文原文
  • 通过强化学习策略指导模型聚焦可疑区域,实现空间感知的推理。
  • 核心级检测达到79.0 AUROC,灵敏度比基线高4.5个百分点。
  • 生成可解释的注意力图,适合临床医生辅助决策与研究使用。

微超声(μUS)是新兴的前列腺癌(PCa)检测成像技术,但病灶识别高度依赖临床经验,导致观察者间差异大。机器学习可降低这种差异,但训练可靠深度模型困难,因标注稀疏且噪声多——通常仅有核心级别的组织病理学结果(如癌症分级和占比),缺乏像素级病变标注,且类别严重不平衡。本文提出Prost-RL,将μUS PCa检测重构为一种空间感知、策略驱动的推理问题,先学习关注何处再进行判断。Prost-RL在基础模型编码器-解码器中嵌入轻量级强化学习策略,生成可解释的空间注意力图,作为癌症概率热图预测和图像级分类的软提示。进一步提出自适应策略优化(APO)以稳定监督-强化学习混合训练,并设计结合对称交叉熵与负熵正则化的抗噪目标,缓解弱标签噪声并促进精准定位。在来自5个临床中心的693名患者共6,607个活检核心的数据集上,Prost-RL实现核心级检测79.0±3.5 AUROC,80%特异性下灵敏度达64.6±6.3%(较最强基线+2.1 AUROC、+4.5灵敏度点),临床显著癌症分类达到79.3±5.8 AUROC。所学策略能突出与活检对齐的区域,提供透明、空间化证据与量化风险预测。代码已开源:https://github.com/DeepRCL/Prost-RL。

原文摘要 · Abstract (English)

Micro-ultrasound ($μ$US) is a new, emerging, and promising imaging modality for prostate cancer (PCa) detection, but accurate identification of suspicious tissue remains highly dependent on clinical experience, leading to substantial inter-observer variability. Machine-learning assistance can reduce this variability; however, training reliable deep models is challenging because supervision is sparse and noisy -- typically limited to core-level histopathology outcomes (e.g., cancer grade and its percentage in a biopsy core) without pixel-level lesion annotations and under severe class imbalance. We introduce Prost-RL, which reframes $μ$US PCa detection as a spatially aware, policy-driven inference problem by learning where to look before decoding. Prost-RL integrates a lightweight reinforcement-learning policy into a foundation-model encoder-decoder to generate interpretable spatial attention maps that act as soft prompts for both cancer-likelihood heatmap prediction and image-level classification. We further propose Adaptive Policy Optimization (APO) to stabilize hybrid supervised-RL training and a noise-robust objective combining symmetric cross-entropy with negative-entropy regularization to mitigate weak-label noise and encourage sharp localization. On a cohort of 6,607 biopsy cores from 693 patients across five clinical sites, Prost-RL achieves $79.0\pm3.5$ AUROC with $64.6\pm6.3$% sensitivity at 80% specificity for core-level detection (+2.1 AUROC and +4.5 sensitivity points over the strongest baseline), and $79.3\pm5.8$ AUROC for clinically significant cancer classification. The learned policy highlights biopsy-aligned regions, providing transparent, spatially grounded evidence alongside quantitative risk predictions. Code is available at: https://github.com/DeepRCL/Prost-RL.

前列腺癌强化学习医学影像可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。