给数据重建攻击定标准,让模型隐私风险可衡量
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
- 提出统一分类与形式化定义,规范攻击研究范式
- 设计量化指标,评估重建质量的精度与多样性
- 用大模型替代人工评估,实现高效高质量测试
数据重建攻击旨在通过有限访问恢复目标模型的训练数据,近年来备受关注。然而,目前尚无关于此类攻击的正式定义和通用评估指标,制约了该领域的发展。本文针对视觉领域提出统一的攻击分类体系与形式化定义,构建一套包含可量化性、一致性、精度和多样性的定量评估指标。同时,利用大语言模型(LLMs)替代人工判断,实现以高质量重建为核心的视觉评估。基于新框架,系统评估现有攻击方法的优劣,并建立未来研究基准。实证结果主要从记忆角度验证了指标有效性,为新型攻击设计提供重要洞见。
原文摘要 · Abstract (English)
Data reconstruction attacks, which aim to recover the training dataset of a target model with limited access, have gained increasing attention in recent years. However, there is currently no consensus on a formal definition of data reconstruction attacks or appropriate evaluation metrics for measuring their quality. This lack of rigorous definitions and universal metrics has hindered further advancement in this field. In this paper, we address this issue in the vision domain by proposing a unified attack taxonomy and formal definitions of data reconstruction attacks. We first propose a set of quantitative evaluation metrics that consider important criteria such as quantifiability, consistency, precision, and diversity. Additionally, we leverage large language models (LLMs) as a substitute for human judgment, enabling visual evaluation with an emphasis on high-quality reconstructions. Using our proposed taxonomy and metrics, we present a unified framework for systematically evaluating the strengths and limitations of existing attacks and establishing a benchmark for future research. Empirical results, primarily from a memorization perspective, not only validate the effectiveness of our metrics but also offer valuable insights for designing new attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。