arXiv:2508.05677cs.CRcs.AI2025-08

研究强化学习问诊系统在对抗攻击下的脆弱性,发现即使有医学约束仍易被攻破。

Adversarial Attacks on Reinforcement Learning-based Medical Questionnaire Systems: Input-level Perturbation Strategies and Medical Constraint Validation

  • 将问诊建模为马尔可夫决策过程,用六种攻击方法生成对抗样本
  • 在NHIS数据集上实现97.6%的临床合理对抗样本生成率,攻击成功率最高达64.70%
  • 适用于医疗AI安全评估、对抗训练与可信医疗系统设计者

基于强化学习的医疗问诊系统在医疗场景中展现出巨大潜力,但其安全性和鲁棒性仍待解决。本研究全面评估了对抗攻击方法,识别并分析其潜在漏洞。将诊断过程建模为马尔可夫决策过程(MDP),状态为患者回答与未提问问题,动作为提问或诊断。实现了六种主流攻击方法:FGSM、PGD、C&W、BIM、DeepFool和AutoAttack,每种使用七个ε值。为确保生成的对抗样本具有临床合理性,构建了包含247条医学约束的验证框架,涵盖生理范围、症状相关性及条件医学约束。在包含182,630个样本的国家健康访谈调查(NHIS)数据集上,针对4年死亡率预测任务,评估了arXiv:2004.00994提出的AdaptiveFS框架。结果表明,对抗攻击显著影响诊断准确率,攻击成功率介于33.08%(FGSM)至64.70%(AutoAttack)之间。研究表明,即使在严格医学输入约束下,此类系统仍存在显著脆弱性。

原文摘要 · Abstract (English)

RL-based medical questionnaire systems have shown great potential in medical scenarios. However, their safety and robustness remain unresolved. This study performs a comprehensive evaluation on adversarial attack methods to identify and analyze their potential vulnerabilities. We formulate the diagnosis process as a Markov Decision Process (MDP), where the state is the patient responses and unasked questions, and the action is either to ask a question or to make a diagnosis. We implemented six prevailing major attack methods, including the Fast Gradient Signed Method (FGSM), Projected Gradient Descent (PGD), Carlini & Wagner Attack (C&W) attack, Basic Iterative Method (BIM), DeepFool, and AutoAttack, with seven epsilon values each. To ensure the generated adversarial examples remain clinically plausible, we developed a comprehensive medical validation framework consisting of 247 medical constraints, including physiological bounds, symptom correlations, and conditional medical constraints. We achieved a 97.6% success rate in generating clinically plausible adversarial samples. We performed our experiment on the National Health Interview Survey (NHIS) dataset (https://www.cdc.gov/nchs/nhis/), which consists of 182,630 samples, to predict the participant's 4-year mortality rate. We evaluated our attacks on the AdaptiveFS framework proposed in arXiv:2004.00994. Our results show that adversarial attacks could significantly impact the diagnostic accuracy, with attack success rates ranging from 33.08% (FGSM) to 64.70% (AutoAttack). Our work has demonstrated that even under strict medical constraints on the input, such RL-based medical questionnaire systems still show significant vulnerabilities.

医疗AI对抗攻击强化学习安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。