arXiv:2410.18057cs.CVcs.CL2024-10ACL被引 40

首个多模态数据删除基准CLEAR,支持文本视觉联合清除敏感信息

CLEAR: Character Unlearning in Textual and Visual Modalities

  • 构建首个公开多模态遗忘基准,含200虚构人物与3700张关联图文
  • 实验证明联合删除文本与图像信息效果优于单一模态方法
  • 适合研究数据隐私、模型可遗忘性及跨模态安全的学者使用

机器遗忘对移除深度学习模型中的隐私或有害信息至关重要。尽管单模态(文本或视觉)遗忘已取得显著进展,但多模态遗忘(MMU)因缺乏评估跨模态数据删除的公开基准而研究不足。为此,我们提出CLEAR,首个专为多模态遗忘设计的开源基准。CLEAR包含200个虚构个体和3700张与其对应的图文问答对,支持跨模态全面评估。我们对11种遗忘方法(如SCRUB、梯度上升、DPO)在四个测试集上进行了综合分析,结果表明联合删除双模态信息的效果优于单模态方法。数据集已发布于https://huggingface.co/datasets/therem/CLEAR。

原文摘要 · Abstract (English)

Machine Unlearning (MU) is critical for removing private or hazardous information from deep learning models. While MU has advanced significantly in unimodal (text or vision) settings, multimodal unlearning (MMU) remains underexplored due to the lack of open benchmarks for evaluating cross-modal data removal. To address this gap, we introduce CLEAR, the first open-source benchmark designed specifically for MMU. CLEAR contains 200 fictitious individuals and 3,700 images linked with corresponding question-answer pairs, enabling a thorough evaluation across modalities. We conduct a comprehensive analysis of 11 MU methods (e.g., SCRUB, gradient ascent, DPO) across four evaluation sets, demonstrating that jointly unlearning both modalities outperforms single-modality approaches. The dataset is available at https://huggingface.co/datasets/therem/CLEAR

机器遗忘多模态数据隐私基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。