arXiv:2602.11661cs.AI2026-02被引 2

提出医疗大模型对齐新范式,兼顾准确、安全与合规。

Quark Medical Alignment: A Holistic Multi-Dimensional Alignment and Collaborative Optimization Paradigm

  • 构建多维对齐矩阵,分解为四大能力维度
  • 实测在真实医疗场景中显著提升对齐效果
  • 适合垂直领域大模型优化,尤其医疗方向

尽管近年来基于人类反馈的强化学习在大语言模型对齐方面进展迅速,但将其迁移至高风险医疗问答任务时暴露出根本性范式不匹配。人类偏好标注成本过高且难以反映医学事实的绝对正确性;可验证奖励强化学习缺乏有效自动验证器,难以处理复杂临床情境。同时,医疗对齐需同时优化准确性、安全性与合规性,但多目标异构奖励信号易引发量纲失配与优化冲突。为此,我们提出一种稳健的医疗对齐范式。首先构建涵盖基础能力、专家知识、在线反馈与格式规范四维度的综合性医疗对齐矩阵,并在每个维度内建立可观测指标→可归因诊断→可优化奖励的闭环,提供细粒度、高分辨率监督信号以支持迭代优化。为解决异构信号导致的梯度主导与优化不稳问题,进一步提出统一优化机制:采用参考冻结归一化对齐奖励量纲,并设计三因子自适应动态加权策略,实现弱项导向、风险优先、冗余减少的协同优化。实验结果表明,该范式在真实医疗场景评估中表现优异,为垂直领域复杂对齐提供了新范式。

原文摘要 · Abstract (English)

While reinforcement learning for large language model alignment has progressed rapidly in recent years, transferring these paradigms to high-stakes medical question answering reveals a fundamental paradigm mismatch. Reinforcement Learning from Human Feedback relies on preference annotations that are prohibitively expensive and often fail to reflect the absolute correctness of medical facts. Reinforcement Learning from Verifiable Rewards lacks effective automatic verifiers and struggles to handle complex clinical contexts. Meanwhile, medical alignment requires the simultaneous optimization of correctness, safety, and compliance, yet multi-objective heterogeneous reward signals are prone to scale mismatch and optimization conflicts. To address these challenges, we propose a robust medical alignment paradigm. We first construct a holistic multi-dimensional medical alignment matrix that decomposes alignment objectives into four categories: fundamental capabilities, expert knowledge, online feedback, and format specifications. Within each category, we establish a closed loop of where observable metrics inform attributable diagnosis, which in turn drives optimizable rewards, thereby providing fine-grained, high-resolution supervision signals for subsequent iterative optimization. To resolve gradient domination and optimization instability problem caused by heterogeneous signals, we further propose a unified optimization mechanism. This mechanism employs Reference-Frozen Normalization to align reward scales and implements a Tri-Factor Adaptive Dynamic Weighting strategy to achieve collaborative optimization that is weakness-oriented, risk-prioritized, and redundancy-reducing. Experimental results demonstrate the effectiveness of our proposed paradigm in real-world medical scenario evaluations, establishing a new paradigm for complex alignment in vertical domains.

医疗AI模型对齐强化学习多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。