用积分反馈机制让医学影像模型更准地发现病灶,避免胡说八道。
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

- 引入积分反馈机制,根据诊断轨迹持续修正奖励信号。
- 在3D CT数据集上异常检测准确率提升12.6%,临床一致性显著增强。
- 适合医疗AI研发者、医学影像分析人员使用,尤其关注诊断可信度。
医学视觉语言模型(VLMs)虽快速发展,但在三维计算机断层扫描(CT)分析中的应用仍受限于优化目标与临床严谨性之间的不匹配。现有强化学习(RL)范式依赖词汇层面的代理信号,导致模型优化语言流畅性而非临床事实,引发诊断级错误。为此,我们提出临床异常评估基底(CABS),将放射科报告分解为可验证的临床语义单元。基于CABS,我们发现标准RL存在“机制性偏差”:表面相似性奖励使策略梯度绕过医学事实。因此,我们提出轨迹积分反馈GRPO(TIF-GRPO),将控制理论融入策略优化。通过将临床推理建模为伪时间轨迹,利用积分反馈回路惩罚持续遗漏(作为累积状态误差)并抑制幻觉(作为过度控制努力)。在3D CT基准测试中,该方法显著提升异常检测与临床一致性,建立医学VLM细粒度调控新范式。项目开源地址:https://github.com/ZJU4HealthCare/TIF-GRPO。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) have rapidly advanced as general-purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT) analysis remains constrained by a persistent mismatch between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms still rely on lexical proxy signals that induce ``\textit{Evaluation Hallucinations}'', where models optimize linguistic fluency rather than factual clinical correctness, leading to diagnostically critical errors. To bridge this gap, we introduce the \textbf{Clinical Abnormality Benchmarking Substrate (CABS)}, a structured system that decomposes radiology reports into verifiable clinical semantic units. Using CABS, we identify a ``\textit{Mechanistic Divergence}'' in standard RL, where surface-similarity rewards drive policy gradients to bypass medical facts. We therefore propose \textbf{Trajectory-Integral Feedback GRPO (TIF-GRPO)}, a novel framework integrating control-theoretic principles into policy optimization. By formulating clinical reasoning as a pseudo-temporal trajectory for anomaly discovery, TIF-GRPO regulates anatomy-aware rewards via an integral feedback loop that penalizes persistent omissions as cumulative state errors and suppresses hallucinations as excessive control effort. Experiments on 3D CT benchmarks demonstrate that our approach significantly enhances abnormality detection and clinical faithfulness, establishing a new paradigm for fine-grained regulation in medical VLMs. Our project is available at \href{https://github.com/ZJU4HealthCare/TIF-GRPO}{GitHub}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。