构建医疗大模型多智能体安全评估框架,发现架构缺陷并提出防御机制。
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
- 设计5000条对抗性医学提示,覆盖25类威胁和100个子主题。
- 共享池架构最脆弱,去中心化架构更具抗干扰能力。
- 提出性格检测与修正机制,恢复系统安全性至基线水平。
随着大语言模型在医疗领域的广泛应用,确保其在多智能体协作配置下的安全性至关重要。本文提出MedSentry,一个包含5000条对抗性医学提示的基准数据集,涵盖25类威胁和100个子主题。结合该数据集,我们构建了端到端的攻防评估流程,系统分析四种典型多智能体拓扑(分层、共享池、集中式、去中心化)在‘暗人格’智能体攻击下的表现。结果揭示各架构在信息污染应对与决策鲁棒性上的关键差异:共享池因开放信息共享而高度脆弱,而去中心化架构凭借冗余与隔离具备更强韧性。为此,我们提出一种基于性格尺度的检测与纠正机制,可识别并修复恶意智能体,使系统安全性能恢复至接近基线水平。MedSentry不仅提供严谨的评估框架,还为医疗领域大模型多智能体系统的安全设计提供实用策略。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed in healthcare, ensuring their safety, particularly within collaborative multi-agent configurations, is paramount. In this paper we introduce MedSentry, a benchmark comprising 5 000 adversarial medical prompts spanning 25 threat categories with 100 subthemes. Coupled with this dataset, we develop an end-to-end attack-defense evaluation pipeline to systematically analyze how four representative multi-agent topologies (Layers, SharedPool, Centralized, and Decentralized) withstand attacks from 'dark-personality' agents. Our findings reveal critical differences in how these architectures handle information contamination and maintain robust decision-making, exposing their underlying vulnerability mechanisms. For instance, SharedPool's open information sharing makes it highly susceptible, whereas Decentralized architectures exhibit greater resilience thanks to inherent redundancy and isolation. To mitigate these risks, we propose a personality-scale detection and correction mechanism that identifies and rehabilitates malicious agents, restoring system safety to near-baseline levels. MedSentry thus furnishes both a rigorous evaluation framework and practical defense strategies that guide the design of safer LLM-based multi-agent systems in medical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。