arXiv:2505.04388cs.CLcs.AI2025-05被引 9

开源医疗大模型Aloe Beta实现高性能与高安全,可媲美闭源顶尖模型。

The Aloe Family Recipe for Open and Specialized Healthcare LLMs

  • 基于Llama 3.1/Qwen 2.5构建,用合成思维链增强公开数据。
  • 通过DPO对齐提升抗越狱攻击能力,多轮评测表现优异。
  • 适合医疗从业者、研究者及关注伦理安全的AI开发者使用。

随着大型语言模型在医疗领域的进步,亟需具备竞争力的开源模型以保障公共利益。本文提出Aloe Family系列模型,优化数据预处理与训练流程,并通过直接偏好优化(DPO)提升模型安全性,结合检索增强生成(RAG)提高有效性。采用包含闭合式、开放式、安全性和人类评估在内的四种测试方法,建立新评价标准。结果表明,Aloe Beta模型在多个医疗基准测试中表现优异,常被医疗专业人员更青睐;在偏见与毒性控制方面显著优于基线,对未见过的越狱攻击具有强鲁棒性。模型以宽松许可协议发布,并附有针对医疗场景的详细风险评估报告。该工作为医疗领域对齐大模型的开发与报告设定了新标杆。

原文摘要 · Abstract (English)

Purpose: With advancements in Large Language Models (LLMs) for healthcare, the need arises for competitive open-source models to protect the public interest. This work contributes to the field of open medical LLMs by optimizing key stages of data preprocessing and training, while showing how to improve model safety (through DPO) and efficacy (through RAG). The evaluation methodology used, which includes four different types of tests, defines a new standard for the field. The resultant models, shown to be competitive with the best private alternatives, are released with a permisive license. Methods: Building on top of strong base models like Llama 3.1 and Qwen 2.5, Aloe Beta uses a custom dataset to enhance public data with synthetic Chain of Thought examples. The models undergo alignment with Direct Preference Optimization, emphasizing ethical and policy-aligned performance in the presence of jailbreaking attacks. Evaluation includes close-ended, open-ended, safety and human assessments, to maximize the reliability of results. Results: Recommendations are made across the entire pipeline, backed by the solid performance of the Aloe Family. These models deliver competitive performance across healthcare benchmarks and medical fields, and are often preferred by healthcare professionals. On bias and toxicity, the Aloe Beta models significantly improve safety, showing resilience to unseen jailbreaking attacks. For a responsible release, a detailed risk assessment specific to healthcare is attached to the Aloe Family models. Conclusion: The Aloe Beta models, and the recipe that leads to them, are a significant contribution to the open-source medical LLM field, offering top-of-the-line performance while maintaining high ethical requirements. This work sets a new standard for developing and reporting aligned LLMs in healthcare.

医疗LLM开源模型安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。