arXiv:2508.04192cs.CV2025-08被引 2

首个生物医学多模态大模型去学习基准,解决隐私泄露与错误知识问题

From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models

  • 构建合成数据生成管道,注入隐私信息和错误知识以模拟真实风险
  • 提出新评估指标,发现现有去学习方法在生物医学场景效果有限
  • 针对医疗隐私与错误知识两类场景,为安全防护提供可测标准

生物医学多模态大语言模型的安全性日益受关注。训练数据常含难以检测的私密信息与错误知识,可能导致部署后泄露隐私或输出错误。直接重训成本过高,机器去学习成为替代方案——通过选择性移除有害样本知识,保留正常能力。然而,当前缺乏评估生物医学多模态大模型去学习效果的数据集。为此,我们提出首个基准测试 MLLMU-Med,基于新型数据生成管道,在训练集中融合合成私密数据与事实性错误。该基准覆盖两大核心场景:1)隐私保护,即患者私密信息误入训练集导致推理时泄露;2)错误知识清除,即来自不可靠来源的错误知识嵌入数据集引发不安全响应。此外,我们提出新的去学习效率评分,综合反映不同子集上的整体表现。在该基准上评估五种去学习方法,结果表明其在移除有害知识方面表现有限,揭示领域仍有巨大提升空间。本工作为该前沿方向建立新研究路径。

原文摘要 · Abstract (English)

The security of biomedical Multimodal Large Language Models (MLLMs) has attracted increasing attention. However, training samples easily contain private information and incorrect knowledge that are difficult to detect, potentially leading to privacy leakage or erroneous outputs after deployment. An intuitive idea is to reprocess the training set to remove unwanted content and retrain the model from scratch. Yet, this is impractical due to significant computational costs, especially for large language models. Machine unlearning has emerged as a solution to this problem, which avoids complete retraining by selectively removing undesired knowledge derived from harmful samples while preserving required capabilities on normal cases. However, there exist no available datasets to evaluate the unlearning quality for security protection in biomedical MLLMs. To bridge this gap, we propose the first benchmark Multimodal Large Language Model Unlearning for BioMedicine (MLLMU-Med) built upon our novel data generation pipeline that effectively integrates synthetic private data and factual errors into the training set. Our benchmark targets two key scenarios: 1) Privacy protection, where patient private information is mistakenly included in the training set, causing models to unintentionally respond with private data during inference; and 2) Incorrectness removal, where wrong knowledge derived from unreliable sources is embedded into the dataset, leading to unsafe model responses. Moreover, we propose a novel Unlearning Efficiency Score that directly reflects the overall unlearning performance across different subsets. We evaluate five unlearning approaches on MLLMU-Med and find that these methods show limited effectiveness in removing harmful knowledge from biomedical MLLMs, indicating significant room for improvement. This work establishes a new pathway for further research in this promising field.

多模态模型安全防护去学习生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。