用中国医德规范训练大模型,让AI做医疗决策更合伦理。
A Human-Centric Pipeline for Aligning Large Language Models with Chinese Medical Ethics
- 构建动态场景评测集,基于260份中医学伦理法规数据。
- 引入专家级评估器,使7B模型在伦理任务上超越更大模型。
- 框架可复用于其他文化法律环境,适合医疗AI安全研究者。
大语言模型在医疗应用中日益广泛,但其与复杂现实场景下医疗伦理的对齐仍缺乏研究。本文提出MedES,一个从260份权威中文医学、伦理和法律文献构建的动态场景评测基准,反映临床决策挑战。为促进模型对齐,设计守护者在环框架,利用基于专家标注数据训练的自动评估器(领域内准确率超97%)生成针对性提示并提供结构化伦理反馈。通过监督微调与领域特定偏好优化,将7B参数模型进行对齐。实验在中文医疗伦理背景下完全开展,结果表明该模型在核心伦理任务上显著优于更大规模基线,在质量和综合评估指标上均有提升。本工作为中文医疗领域大模型伦理对齐提供了可落地、可扩展的框架,并提示类似管道可通过替换规范语料库推广至其他法律文化环境。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled their application to a range of healthcare tasks. However, aligning LLMs with the nuanced demands of medical ethics, especially under complex real world scenarios, remains underexplored. In this work, we present MedES, a dynamic, scenario-centric benchmark specifically constructed from 260 authoritative Chinese medical, ethical, and legal sources to reflect the challenges in clinical decision-making. To facilitate model alignment, we introduce a guardian-in-the-loop framework that leverages a dedicated automated evaluator (trained on expert-labeled data and achieving over 97% accuracy within our domain) to generate targeted prompts and provide structured ethical feedback. Using this pipeline, we align a 7B-parameter LLM through supervised fine-tuning and domain-specific preference optimization. Experimental results, conducted entirely within the Chinese medical ethics context, demonstrate that our aligned model outperforms notably larger baselines on core ethical tasks, with observed improvements in both quality and composite evaluation metrics. Our work offers a practical and adaptable framework for aligning LLMs with medical ethics in the Chinese healthcare domain, and suggests that similar alignment pipelines may be instantiated in other legal and cultural environments through modular replacement of the underlying normative corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。