提升医学影像问答的逻辑一致性,让AI推理更可靠。
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
- 设计新奖励机制,确保视觉理解与推理一致
- 在30种医学影像上实现跨任务强泛化性能
- 适合医疗AI研发者和临床辅助系统开发者
在医学视觉问答中,准确回答依赖于三个关键步骤:精准感知医学影像数据、基于视觉输入与文本问题的逻辑推理,以及从推理过程生成连贯答案。当前通用视觉语言模型通过大规模强化学习显著提升推理能力,但在医疗领域受限于两大挑战:感知与推理阶段的不一致,以及推理路径与答案生成间的波动性,且缺乏高质量医疗数据支持大规模强化学习。本文提出Med-Zero-17K,一个专为纯强化学习训练构建的标注数据集,涵盖超过30种医学影像模态和24项临床任务。进一步提出一致性感知偏好优化框架CAPO,通过奖励机制保障感知-推理间的忠实度、推理到答案的一致性,以及最终输出的规则准确性。在域内与域外场景的大量实验表明,该方法优于多个强基线模型,在3D医学视觉问答基准和R1类训练范式下均展现卓越泛化能力。
原文摘要 · Abstract (English)
In medical visual question answering (Med-VQA), achieving accurate responses relies on three critical steps: precise perception of medical imaging data, logical reasoning grounded in visual input and textual questions, and coherent answer derivation from the reasoning process. Recent advances in general vision-language models (VLMs) show that large-scale reinforcement learning (RL) could significantly enhance both reasoning capabilities and overall model performance. However, their application in medical domains is hindered by two fundamental challenges: 1) misalignment between perceptual understanding and reasoning stages, and 2) inconsistency between reasoning pathways and answer generation, both compounded by the scarcity of high-quality medical datasets for effective large-scale RL. In this paper, we first introduce Med-Zero-17K, a curated dataset for pure RL-based training, encompassing over 30 medical image modalities and 24 clinical tasks. Moreover, we propose a novel large-scale RL framework for Med-VLMs, Consistency-Aware Preference Optimization (CAPO), which integrates rewards to ensure fidelity between perception and reasoning, consistency in reasoning-to-answer derivation, and rule-based accuracy for final responses. Extensive experiments on both in-domain and out-of-domain scenarios demonstrate the superiority of our method over strong VLM baselines, showcasing strong generalization capability to 3D Med-VQA benchmarks and R1-like training paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。