arXiv:2602.00400cs.AI2026-02

用知识引导的强化学习,让医疗问答模型推理更稳定、更准确。

KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA

  • 只对高质量推理路径进行教师指导,避免错误误导
  • 在医学问答任务中,泛化性能超越传统强化学习方法
  • 适合需要可靠推理的医疗AI应用

强化学习(RL)已成为诱导大语言模型和视觉-语言模型显式推理行为的有前景范式。然而,由于轨迹级奖励稀疏,基于推理的强化学习后训练仍面临根本性挑战,导致信用分配模糊和严重探索失败,使策略陷入‘学习悬崖’。近期的在线蒸馏方法引入密集教师监督以稳定优化,但对所有生成轨迹统一应用。我们认为,这种均匀蒸馏不适用于高推理强度任务,因为低质量轨迹常源于早期逻辑错误,错误情境下的蒸馏会注入噪声且错位的梯度。为此,我们提出知识增强型偏好优化(KEPO),一个统一的后训练框架,包含:(i) 质量门控的在线蒸馏目标,仅对高质量轨迹施加密集教师指导;(ii) 知识增强的探索策略,利用教师模型学习到的提示,有选择地采样奖励正向的在线轨迹用于强化学习,从而缓解探索崩溃。在单源泛化设置下的医学视觉问答基准上评估,KEPO展现出更高的训练稳定性、更连贯的推理行为,以及优于强化学习与在线蒸馏基线的分布外性能。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as a promising paradigm for inducing explicit reasoning behaviors in large language and vision-language models. However, reasoning-oriented RL post-training remains fundamentally challenging due to sparse trajectory-level rewards, leading to ambiguous credit assignment and severe exploration failures that can trap the policy in a ``learning cliff.'' Recent on-policy distillation methods introduce dense teacher supervision to stabilize optimization, but apply it uniformly across all generated trajectories. We argue that such uniform distillation is ill-suited for reasoning-intensive tasks, as low-quality on-policy trajectories often originate from early logical errors, and distillation under flawed contexts injects noisy and misaligned gradients. To address these challenges, we propose Knowledge-Enhanced Preference Optimization (KEPO), a unified post-training framework that integrates: (i) a quality-gated on-policy distillation objective that selectively applies dense teacher guidance only to high-quality trajectories, and (ii) a knowledge-enhanced exploration strategy that leverages hints learned from a teacher model to rejectively sample reward-positive on-policy trajectories for RL, thereby mitigating exploration collapse. Evaluated on a challenging medical visual question answering benchmark under single-source generalization, KEPO demonstrates improved training stability, more coherent reasoning behaviors, and superior out-of-distribution performance over reinforcement learning and on-policy distillation baselines.

医疗问答强化学习多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。