用多智能体循环评估提升医疗大模型安全性和可靠性。
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops
- 设计多智能体迭代对齐框架,结合生成与评估模型协同优化。
- 伦理违规减少89%,风险等级下降92%,收敛更快。
- 适合关注医疗AI合规与安全的开发者和监管机构。
大型语言模型在医疗领域应用日益广泛,但其伦理完整性和安全性仍是临床部署的主要障碍。本文提出一种多智能体精炼框架,通过结构化、迭代式的对齐机制提升医疗大模型的安全性与可靠性。系统融合DeepSeek R1与Med-PaLM两个生成模型,以及LLaMA 3.1和Phi-4两个评估智能体,分别依据美国医学会(AMA)医学伦理原则与五级安全风险评估(SRA-5)协议进行响应评价。在涵盖九个伦理领域的900个临床多样性查询上评估性能,衡量收敛效率、伦理违规减少及特定领域风险行为。结果表明,DeepSeek R1平均仅需2.34次迭代即实现收敛(低于Med-PaLM的2.67次),而Med-PaLM在隐私敏感场景中表现更优。该迭代多智能体循环使伦理违规减少89%,风险等级下降92%,验证了方法的有效性。本研究提出了一个可扩展、符合监管要求且成本可控的医疗AI安全管理范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly applied in healthcare, yet ensuring their ethical integrity and safety compliance remains a major barrier to clinical deployment. This work introduces a multi-agent refinement framework designed to enhance the safety and reliability of medical LLMs through structured, iterative alignment. Our system combines two generative models - DeepSeek R1 and Med-PaLM - with two evaluation agents, LLaMA 3.1 and Phi-4, which assess responses using the American Medical Association's (AMA) Principles of Medical Ethics and a five-tier Safety Risk Assessment (SRA-5) protocol. We evaluate performance across 900 clinically diverse queries spanning nine ethical domains, measuring convergence efficiency, ethical violation reduction, and domain-specific risk behavior. Results demonstrate that DeepSeek R1 achieves faster convergence (mean 2.34 vs. 2.67 iterations), while Med-PaLM shows superior handling of privacy-sensitive scenarios. The iterative multi-agent loop achieved an 89% reduction in ethical violations and a 92% risk downgrade rate, underscoring the effectiveness of our approach. This study presents a scalable, regulator-aligned, and cost-efficient paradigm for governing medical AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。