arXiv:2511.07481cs.LG2025-11被引 3

对比微调前后大模型在基因数据上的隐私泄露风险,发现微调能提升抗攻击能力。

Comparing Reconstruction Attacks on Pretrained Versus Full Fine-tuned Large Language Model Embeddings on Homo Sapiens Splice Sites Genomic Data

  • 针对基因序列设计专用分词器,比较预训练与微调模型的嵌入表示。
  • 微调后模型抗重构攻击能力提升,XLNet、GPT-2、BERT分别提高19.8%、9.8%、7.8%。
  • 适合关注基因组数据隐私保护的研究者和开发者参考。

本研究探讨大语言模型在基因序列上的嵌入重构攻击问题,重点分析微调对隐私脆弱性的影响。基于Pan等人的开创性工作,我们利用HS3D基因组数据集,系统评估预训练与全微调模型嵌入在重构攻击下的表现差异。研究拓展了前人工作的三个维度:首先,将重构攻击流程应用于两类嵌入,填补方法学空白;其次,设计专用于DNA序列的分词机制,提升模型对基因数据的处理能力;第三,从位置特异性、核苷酸类型和隐私变化角度进行细致对比分析。结果表明,微调显著降低嵌入的重构风险——在XLNet、GPT-2、BERT上分别提升19.8%、9.8%、7.8%,说明任务优化可成为潜在的隐私增强手段。研究强调处理敏感基因数据时需部署高级防护机制,并指出微调具有进一步探索价值。

原文摘要 · Abstract (English)

This study investigates embedding reconstruction attacks in large language models (LLMs) applied to genomic sequences, with a specific focus on how fine-tuning affects vulnerability to these attacks. Building upon Pan et al.'s seminal work demonstrating that embeddings from pretrained language models can leak sensitive information, we conduct a comprehensive analysis using the HS3D genomic dataset to determine whether task-specific optimization strengthens or weakens privacy protections. Our research extends Pan et al.'s work in three significant dimensions. First, we apply their reconstruction attack pipeline to pretrained and fine-tuned model embeddings, addressing a critical gap in their methodology that did not specify embedding types. Second, we implement specialized tokenization mechanisms tailored specifically for DNA sequences, enhancing the model's ability to process genomic data, as these models are pretrained on natural language and not DNA. Third, we perform a detailed comparative analysis examining position-specific, nucleotide-type, and privacy changes between pretrained and fine-tuned embeddings. We assess embeddings vulnerabilities across different types and dimensions, providing deeper insights into how task adaptation shifts privacy risks throughout genomic sequences. Our findings show a clear distinction in reconstruction vulnerability between pretrained and fine-tuned embeddings. Notably, fine-tuning strengthens resistance to reconstruction attacks in multiple architectures -- XLNet (+19.8\%), GPT-2 (+9.8\%), and BERT (+7.8\%) -- pointing to task-specific optimization as a potential privacy enhancement mechanism. These results highlight the need for advanced protective mechanisms for language models processing sensitive genomic data, while highlighting fine-tuning as a potential privacy-enhancing technique worth further exploration.

基因组隐私安全大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。