LoRA-MINT可精准检测微调大模型是否用过特定数据,保护数据版权与隐私。
Auditing Training Data in Domain-adapted LLMs: LoRA-MINT

- 基于困惑度分析设计新型成员推断方法,适配低秩微调的LLM。
- 在3个数据集上精度达0.77~0.92,优于现有方法。
- 适用于各类微调模型,助力AI伦理与数据合规审计。
我们提出LoRA-MINT,一种针对通过低秩适应(LoRA)微调的大型语言模型(LLM)的成员推断测试(MINT)新方法。其主要目标是评估单个样本是否曾被用于这些适配模型的训练数据,为知识产权和敏感数据管理提供有效审计工具。我们分析了模型困惑度与成员身份之间的关系,建立系统化框架以估算微调后LLM的数据暴露程度。在四个模型和三个基准数据集上进行实验,判定给定数据是否用于训练的精度范围为0.77至0.92,显著优于当前最先进基线,验证了该方法的鲁棒性与通用性。总体而言,研究结果强调了LoRA-MINT作为高效、可扩展的LLM审计框架的潜力,有助于提升透明度,推动人工智能与自然语言处理技术的伦理化与负责任部署。尽管本文讨论和实验聚焦于LoRA调整的LLM,但所提方法多数可推广至其他微调技术或更广泛的领域适配型AI模型。
原文摘要 · Abstract (English)
We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Natural Language Processing (NLP) tasks through Low-Rank Adaptation (LoRA). The primary goal is to assess whether individual samples were part of the training data of these adapted models, providing a useful auditing tool for the management of intellectual property and sensitive data. Our analysis explores the relationship between model perplexity and membership status, providing a systematic framework for estimating data exposure in fine-tuned LLMs. We conducted experiments on four models and three benchmark datasets, obtaining precision values in determining if given data were used for training ranging from 0.77 to 0.92, which outperform state-of-the-art baselines and demonstrate the robustness and generality of the proposed method. In general, our findings underscore the potential of LoRA-MINT as an effective and scalable framework for auditing LLMs, improving transparency, and fostering the ethical and responsible deployment of AI and NLP technologies. For the sake of concreteness and current relevance, our discussion and experiments are centered on LoRAadjusted LLMs, but note that most of the presented methodology is easily applicable for auditing training data given any other technique for adapting LLMs or, more generally, any other domain-adapted AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。