arXiv:2606.25407cs.CV2026-06

用教师自进化与竞赛机制,提升医学影像问答的推理质量。

Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher

论文配图:Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher
图 1 · 摘自论文原文
  • 教师与推理器对抗迭代,通过比较监督优化思维链
  • 在多个胸部X光问答数据集上超越强基线模型
  • 适合需要可解释医疗AI的临床辅助系统研究者

胸部X光视觉问答(CXR VQA)要求模型不仅给出正确答案,还需生成可靠的医学推理。现有基于强化学习的方法通常依赖答案级奖励,这类信号过于粗略,难以提升思维链(CoT)质量,且当群体优势归零时会失效。本文提出教-思框架(Teach-to-Reason, T2R),引入基于比较的监督:一个自我进化的教师生成参考推理,推理器则与之竞争优化。随着教师不断强化,推理器面对的参考也逐步增强。我们设计了案例级奖励机制,在原始奖励有效时保留其正负划分,在奖励退化时转而使用竞争得分恢复监督。在多个开放问答基准上的实验表明,T2R持续优于强基线模型,证明在受控且合理的设计下,基于比较的监督能更有效地优化推理过程。

原文摘要 · Abstract (English)

Chest X-ray visual question answering (CXR VQA) requires models not only to predict correct answers, but also to produce reliable medical reasoning. However, existing reinforcement-learning-based training typically relies on answer-level rewards, which are often too coarse to improve chain-of-thought (CoT) quality and can become ineffective when group-level advantages collapse to zero. We propose \textbf{Teach-to-Reason (T2R)}, a framework that introduces comparison-based supervision into CoT optimization through a self-improving \emph{Teacher} and a competition-guided \emph{Reasoner}. As the Teacher is iteratively strengthened via self-competition, the Reasoner is optimized against progressively stronger Teacher-generated references. We further introduce a case-wise reward design that preserves the original reward-induced positive/negative partition when it is informative, and restores supervision from competition scores when the original reward signal degenerates. Experiments on multiple CXR open-ended VQA benchmarks show that T2R consistently outperforms strong baselines, indicating that comparison-based supervision, when integrated in a controlled and principled manner, provides a more effective training signal for reasoning optimization.

医学AI视觉问答推理优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。