用医生验证的医学标准提升大模型临床输出质量
ClinAlign: Scaling Healthcare Alignment from Clinician Preference
- 构建医生验证的健康评分标准数据集,实现精准偏好对齐
- 提炼出119条可复用的临床原则,支持高效监督
- 在小模型上达到领先性能,适合医疗AI落地应用
尽管大型语言模型具备专家级医学知识,但将其开放式输出与细致的临床偏好对齐仍具挑战。现有方法多依赖粗粒度目标或缺乏专业依据的自动化评判。本文提出两阶段框架:首先构建包含7,034个医生验证的偏好样本的HealthRubrics数据集,由临床医生修正LLM生成的评分标准以符合严格医学规范;其次将这些标准提炼为119条广泛适用、基于临床维度的HealthPrinciples,实现可扩展的监督。利用HealthPrinciples,我们实现离线对齐(合成未标注查询的评分标准)和推理时引导自修正。使用该框架训练的30B-A3B模型在HealthBench-Hard上取得33.4%得分,超越更大型模型如Deepseek-R1和o3,建立资源高效临床对齐基准。
原文摘要 · Abstract (English)
Although large language models (LLMs) demonstrate expert-level medical knowledge, aligning their open-ended outputs with fine-grained clinician preferences remains challenging. Existing methods often rely on coarse objectives or unreliable automated judges that are weakly grounded in professional guidelines. We propose a two-stage framework to address this gap. First, we introduce HealthRubrics, a dataset of 7,034 physician-verified preference examples in which clinicians refine LLM-drafted rubrics to meet rigorous medical standards. Second, we distill these rubrics into HealthPrinciples: 119 broadly reusable, clinically grounded principles organized by clinical dimensions, enabling scalable supervision beyond manual annotation. We use HealthPrinciples for (1) offline alignment by synthesizing rubrics for unlabeled queries and (2) an inference-time tool for guided self-revision. A 30B-A3B model trained with our framework achieves 33.4% on HealthBench-Hard, outperforming much larger models including Deepseek-R1 and o3, establishing a resource-efficient baseline for clinical alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。