医学推理模型答对题但格式一变就错,需评估其格式鲁棒性。
On the Robustness of Answer Formats in Medical Reasoning Models
- 测试模型在多选、问答、排序三种答案格式下的表现
- 15个模型格式鲁棒性差异大,35%-100%不等
- 监督微调比强化学习更稳定,奖励设计影响关键
医学推理模型(MRMs)在医疗基准上表现优于医学大模型;然而,高准确率不足以支撑实际应用。真实场景中一个重要要求是应对不同输出格式的鲁棒性:同一问题若要求不同答案格式,结果正确性不应改变。本文聚焦于该现象,提出“答案格式鲁棒性”指标,衡量模型在多种格式下生成正确输出的一致性。研究涵盖三种典型格式:多选题、开放式问答、排序列表。在15个专有及开源模型中,观察到格式鲁棒性显著差异(35%-100%)。进一步在共享主干模型上进行受控微调实验,使用相同训练数据隔离微调范式影响。结果表明,监督微调使模型跨格式行为更稳定,而强化学习微调常表现出更高脆性,且不稳定性程度强烈依赖于奖励设计。总体而言,医学推理模型的答案格式鲁棒性可训练但易碎,实际应用前需谨慎评估。
原文摘要 · Abstract (English)
Medical reasoning models (MRMs) achieve superior performance on medical benchmarks compared to medical LLMs; however, high accuracy alone is insufficient for practical deployment. One of such requirements for real-world application is robustness to varying output constraints. Specifically, posing the same medical question while requesting different answer formats should not affect the underlying correctness of the response. We investigate this phenomenon in this paper, focusing on MRMs. To quantify this behavior, we propose the metric answer-format robustness: the ability to reliably generate correct outputs across varying specified formats. We examine three representative formats: multiple-choice, open-ended question-answering, and ranked lists. Across 15 proprietary and open-weight models, we observe substantial variation in format robustness (35-100%). Furthermore, we conduct controlled fine-tuning experiments on a shared backbone with matched training data to isolate the effects of the fine-tuning paradigm. We find that supervised fine-tuning yields more stable behavior across formats, whereas reinforcement fine-tuning often exhibits higher cross-format brittleness, with the degree of instability strongly dependent on reward design. Overall, answer-format robustness in MRMs is trainable yet brittle and requires careful evaluation for practical medical use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。