arXiv:2605.08257cs.CRcs.AI2026-05中稿 · oral presentation …

提出医疗决策智能体安全增强框架,显著提升对抗攻击下的可靠性。

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

  • 构建六阶段协同安全框架,融合风险感知与知识一致性验证。
  • 在多种攻击下攻击成功率降至8.7%,知识一致性达0.91。
  • 模块消融实验证明各环节贡献,适合医疗AI安全研究者参考。

为提升医疗决策智能体的对抗鲁棒性、安全性和可信度,本文提出全链路安全增强框架,涵盖输入风险感知、医学证据约束、知识一致性验证、决策置信度重加权、安全输出控制及对抗反馈更新。设计ARSM-Agent,采用加权联合目标函数,包含决策准确率损失、对抗鲁棒性损失、安全拒答损失和知识一致性损失,权重分别为0.3、0.3、0.2、0.2。通过多模块协同实现完整医疗决策流程。实验表明,该算法优于四种基线模型(LLM-Agent、Retrieval-Agent、Filter-Agent、Adv-Train-Agent)。在语义扰动、提示注入、药物名混淆及虚假证据攻击下,整体攻击成功率降至8.7%,知识一致性得分达0.91。消融实验量化各模块贡献:移除风险感知、证据检索、一致性验证和置信度重加权分别导致准确率下降6.7%、9.1%、7.6%、4.4%,攻击成功率上升13.8%、11.1%、8.6%、6.9%。该方法有效解决医疗决策智能体的关键安全问题,在复杂场景中实现可靠决策,为医疗AI提供可信智能支持。

原文摘要 · Abstract (English)

Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develops a full-link security enhancement framework, which describes "input risk perception - medical evidence constraint - knowledge consistency verification - decision confidence reweighting - security output control - adversarial feedback update." We propose ARSM-Agent and define a weighted joint objective consisting of decision accuracy loss, adversarial robustness loss, safety refusal loss, and knowledge consistency loss, with weights of 0.3, 0.3, 0.2, and 0.2, respectively. The whole medical decision formulation is implemented by multi-module collaborative linkage. We verify that the algorithm is more efficient than four baselines, including LLM-Agent, Retrieval-Agent, Filter-Agent, and Adv-Train-Agent. Under semantic perturbation, prompt injection, drug-name confusion, and false-evidence attacks, ARSM-Agent reduces the overall attack success rate to 8.7% and achieves a knowledge consistency score of 0.91. Ablation experiments quantify each module's contribution: removing risk perception, evidence retrieval, consistency verification, and confidence reweighting reduces accuracy by 6.7%, 9.1%, 7.6%, and 4.4%, respectively, and increases attack success rate by 13.8%, 11.1%, 8.6%, and 6.9%. The proposed approach addresses key security issues of medical decision making intelligent agents, obtains secure decision making in challenging scenarios, and provides reliable intelligent support for medical decision-making intelligent agents.

医疗AI对抗鲁棒智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。