评估基因组语言模型的记忆风险,发现其存在可量化的隐私泄露隐患。
Quantifying Memorization and Privacy Risks in Genomic Language Models
- 构建多维度隐私评估框架,融合困惑度、特洛伊序列提取与成员推断
- 实验证明模型会记忆训练数据中的重复序列,风险随模型容量上升
- 强调需对基因组AI系统进行多向隐私审计,避免单一攻击方式遗漏风险
基因组语言模型(GLMs)在变异预测、调控元件识别和跨任务迁移学习中展现出强大能力。然而,随着这些模型在敏感基因组数据上训练或微调,其可能记忆具体序列,引发隐私泄露与合规风险。现有研究多关注通用语言模型的记忆风险,但对具有固定碱基字母表、强生物结构与个体可识别性的基因组数据缺乏系统性评估。本文提出一个综合性的多向隐私评估框架,整合困惑度检测、特洛伊序列提取与成员推断三种方法,形成统一的最坏情况记忆风险评分。通过在合成与真实基因组数据中植入不同重复率的特洛伊序列,精确量化重复频率与训练动态对记忆的影响。我们在多种GLM架构上评估该框架,发现模型确实存在可测量的记忆现象,且风险程度随模型容量与训练策略而异。结果表明,单一攻击手段无法覆盖全部记忆风险,强调必须将多向隐私审计作为基因组人工智能系统的标准实践。
原文摘要 · Abstract (English)
Genomic language models (GLMs) have emerged as powerful tools for learning representations of DNA sequences, enabling advances in variant prediction, regulatory element identification, and cross-task transfer learning. However, as these models are increasingly trained or fine-tuned on sensitive genomic cohorts, they risk memorizing specific sequences from their training data, raising serious concerns around privacy, data leakage, and regulatory compliance. Despite growing awareness of memorization risks in general-purpose language models, little systematic evaluation exists for these risks in the genomic domain, where data exhibit unique properties such as a fixed nucleotide alphabet, strong biological structure, and individual identifiability. We present a comprehensive, multi-vector privacy evaluation framework designed to quantify memorization risks in GLMs. Our approach integrates three complementary risk assessment methodologies: perplexity-based detection, canary sequence extraction, and membership inference. These are combined into a unified evaluation pipeline that produces a worst-case memorization risk score. To enable controlled evaluation, we plant canary sequences at varying repetition rates into both synthetic and real genomic datasets, allowing precise quantification of how repetition and training dynamics influence memorization. We evaluate our framework across multiple GLM architectures, examining the relationship between sequence repetition, model capacity, and memorization risk. Our results establish that GLMs exhibit measurable memorization and that the degree of memorization varies across architectures and training regimes. These findings reveal that no single attack vector captures the full scope of memorization risk, underscoring the need for multi-vector privacy auditing as a standard practice for genomic AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。