给蛋白质生成模型加数字水印,防抄袭还防滥用。
FoldMark: Protecting Protein Generative Models with Watermarking
- 用两阶段方法在蛋白结构中嵌入用户专属水印
- 水印几乎不影响结构质量,恢复率超95%
- 适合需版权保护的生物设计与药物研发场景
蛋白质结构是理解其功能的关键,对生物工程、药物发现和分子生物学至关重要。近年来,生成式AI显著提升了计算蛋白质结构预测与设计的性能。然而,版权保护与有害内容生成(生物安全)等伦理问题制约了其广泛应用。本文探讨在蛋白质生成模型及其输出中嵌入水印的可能性,以实现版权认证与生成结构追踪。作为概念验证,提出一种名为FoldMark的通用水印策略:首先预训练水印编码器与解码器,可微调蛋白结构以嵌入用户信息并准确恢复;其次,通过水印条件下的低秩适配(LoRA)微调生成模型,在保持生成质量的同时学习生成高恢复率的水印结构。在ESMFold、MultiFlow、FrameDiff和FoldFlow等开源模型上进行大量实验,结果表明该方法对所有生成模型均有效。同时,水印框架对原始结构质量影响极小,且对后处理和自适应攻击具有鲁棒性。
原文摘要 · Abstract (English)
Protein structure is key to understanding protein function and is essential for progress in bioengineering, drug discovery, and molecular biology. Recently, with the incorporation of generative AI, the power and accuracy of computational protein structure prediction/design have been improved significantly. However, ethical concerns such as copyright protection and harmful content generation (biosecurity) pose challenges to the wide implementation of protein generative models. Here, we investigate whether it is possible to embed watermarks into protein generative models and their outputs for copyright authentication and the tracking of generated structures. As a proof of concept, we propose a two-stage method FoldMark as a generalized watermarking strategy for protein generative models. FoldMark first pretrain watermark encoder and decoder, which can minorly adjust protein structures to embed user-specific information and faithfully recover the information from the encoded structure. In the second step, protein generative models are fine-tuned with watermark-conditioned Low-Rank Adaptation (LoRA) modules to preserve generation quality while learning to generate watermarked structures with high recovery rates. Extensive experiments are conducted on open-source protein structure prediction models (e.g., ESMFold and MultiFlow) and de novo structure design models (e.g., FrameDiff and FoldFlow) and we demonstrate that our method is effective across all these generative models. Meanwhile, our watermarking framework only exerts a negligible impact on the original protein structure quality and is robust under potential post-processing and adaptive attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。