arXiv:2511.18084cs.LGcs.AI2025-11被引 1

医学大模型越精准,医生越不信任,因解释不清且难落地。

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement from Clinical Decision-making Quality

  • 用8000+不孕治疗数据对比四种对齐方法,发现强化学习最优但医生不选。
  • 医生更偏爱微调模型,因其推理清晰、治疗方案可行性强。
  • 模型算法越准,临床采纳率越低,凸显技术与医疗需求的脱节。

大型语言模型在临床决策支持中应用日益广泛,但其与真实医学复杂推理路径的对齐仍是重大挑战。基于8000余份不孕治疗记录,我们通过双层评估框架(自动基准+盲评医生参与)系统比较了四种对齐策略:监督微调(SFT)、直接偏好优化(DPO)、组相对策略优化(GRPO)和上下文学习(ICL)。GRPO在多层决策任务中达到最高算法准确率,验证了强化学习在结构化预测中的优势。然而,医生一致偏好SFT模型,认为其推理过程更清晰(p = 0.035),治疗可行性更高(p = 0.019)。在盲态配对比较中,SFT胜出率高达51.2%,高于GRPO(26.2%)和医生原决策(22.7%)。结果揭示了‘对齐悖论’:算法性能提升未必带来临床信任,甚至可能偏离以人为本的需求。研究强调,对齐策略应优先考虑临床可解释性与实际可行性,而非仅追求决策准确率。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly adopted in clinical decision support, yet aligning them with the multifaceted reasoning pathways of real-world medicine remains a major challenge. Using more than 8,000 infertility treatment records, we systematically evaluate four alignment strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and In-Context Learning (ICL) through a dual-layer framework combining automatic benchmarks with blinded doctor-in-the-loop assessments. GRPO achieves the highest algorithmic accuracy across multiple decision layers, confirming the value of reinforcement-based optimization for structured prediction tasks. However, clinicians consistently prefer the SFT model, citing clearer reasoning processes (p = 0.035) and higher therapeutic feasibility (p = 0.019). In blinded pairwise comparisons, SFT attains the highest winning rate (51.2%), outperforming both GRPO (26.2%) and even physicians' original decisions (22.7%). These results reveal an alignment paradox: algorithmic improvements do not necessarily translate into higher clinical trust, and may diverge from human-centered preferences. Our findings highlight the need for alignment strategies that prioritize clinically interpretable and practically feasible reasoning, rather than solely optimizing decision-level accuracy.

医疗AI大模型对齐临床决策可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。