arXiv:2506.01257cs.CLcs.AI2025-06综述被引 14

开源医疗大模型DeepSeek-R1在诊断与推理上媲美闭源模型,但存在安全风险。

DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models

  • 融合专家混合与思维链,提升医疗与数学等领域的推理能力。
  • 在USMLE和AIME测试中表现优异,儿科与眼科决策支持效果突出。
  • 适合资源受限场景部署,但需警惕偏见与误导性输出风险。

DeepSeek-R1是由DeepSeek开发的前沿开源大语言模型,采用混合专家(MoE)、思维链(CoT)推理与强化学习相结合的混合架构,在数学、医疗诊断、代码生成和药物研究等结构化问题求解领域展现出卓越能力。该模型以宽松的MIT许可证发布,为GPT-4o和Claude-3 Opus等闭源模型提供了透明且低成本的替代方案。其在美国医学执照考试(USMLE)和美国邀请数学竞赛(AIME)等基准测试中表现优异,尤其在儿科和眼科临床决策支持任务中成果显著。模型具备高效推理能力,同时保持深层推理性能,适用于资源受限环境。然而,该模型在多语言及伦理敏感场景下更易受偏见、虚假信息、对抗攻击和安全失效影响。本文综述了其在可解释性、可扩展性与适应性方面的优势,也指出其在自然语言流畅性与安全对齐上的不足。未来研究应聚焦偏见缓解、语言理解优化、领域验证与合规监管。整体而言,DeepSeek-R1推动了开放、可扩展AI的发展,亟需协同治理以保障其负责任、公平的部署。

原文摘要 · Abstract (English)

DeepSeek-R1 is a cutting-edge open-source large language model (LLM) developed by DeepSeek, showcasing advanced reasoning capabilities through a hybrid architecture that integrates mixture of experts (MoE), chain of thought (CoT) reasoning, and reinforcement learning. Released under the permissive MIT license, DeepSeek-R1 offers a transparent and cost-effective alternative to proprietary models like GPT-4o and Claude-3 Opus; it excels in structured problem-solving domains such as mathematics, healthcare diagnostics, code generation, and pharmaceutical research. The model demonstrates competitive performance on benchmarks like the United States Medical Licensing Examination (USMLE) and American Invitational Mathematics Examination (AIME), with strong results in pediatric and ophthalmologic clinical decision support tasks. Its architecture enables efficient inference while preserving reasoning depth, making it suitable for deployment in resource-constrained settings. However, DeepSeek-R1 also exhibits increased vulnerability to bias, misinformation, adversarial manipulation, and safety failures - especially in multilingual and ethically sensitive contexts. This survey highlights the model's strengths, including interpretability, scalability, and adaptability, alongside its limitations in general language fluency and safety alignment. Future research priorities include improving bias mitigation, natural language comprehension, domain-specific validation, and regulatory compliance. Overall, DeepSeek-R1 represents a major advance in open, scalable AI, underscoring the need for collaborative governance to ensure responsible and equitable deployment.

大模型医疗AI开源模型安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。