arXiv:2603.02262cs.CRcs.AI2026-03被引 1

在微调阶段悄悄污染医疗大模型的推理过程,让其在特定领域表现骤降。

Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs

论文配图:Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs
图 1 · 摘自论文原文
  • 通过注入错误推理链污染少量训练数据,隐蔽攻击模型推理能力。
  • 无正确样本时,目标领域准确率显著下降,且需最少数量和比例的毒化样本。
  • 比灾难性遗忘更高效精准,揭示医疗大模型微调阶段的安全风险。

监督微调(SFT)对医疗大语言模型(LLM)的发展至关重要,但以往的中毒研究主要关注可检测的后门攻击。本文提出一种针对医疗LLM推理过程的新中毒攻击方法。与后门攻击不同,该方法在少样本训练数据中注入毒化推理链,导致模型在特定医疗主题上性能悄然退化。实验表明,知识覆盖无效,而推理链中毒在训练集中无正确样本时,显著降低目标主题的准确率。有效且隐蔽的攻击需最小数量和比例的毒化样本,其效率与准确性优于灾难性遗忘。本研究揭示了SFT阶段中毒的风险,呼吁在敏感医疗领域加强防御研究。

原文摘要 · Abstract (English)

Supervised fine-tuning (SFT) is essential for the development of medical large language models (LLMs), yet prior poisoning studies have mainly focused on the detectable backdoor attacks. We propose a novel poisoning attack targeting the reasoning process of medical LLMs during SFT. Unlike backdoor attacks, our method injects poisoned rationales into few-shot training data, leading to stealthy degradation of model performance on targeted medical topics. Results showed that knowledge overwriting was ineffective, while rationale poisoning caused significant decline on the accuracy of the target subject, as long as no correct samples of the same subject appear in the dataset. A minimum number and ratio of poisoned samples was needed to carry out an effective and stealthy attack, which was more efficient and accurate than catastrophic forgetting. We demonstrate though this study the risk of SFT-stage poisoning, hoping to spur more studies of defense in the sensitive medical domain.

医疗AI模型安全中毒攻击微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。