用推理时控制避免蛋白模型生成有毒序列
Inference-Time Toxicity Mitigation in Protein Language Models
- 通过放大基线与毒性微调模型的logit差异,实现无需重训练的推理控制
- 在4个分类群中将预测毒性率降至微调基线以下,且保持蛋白结构合理性
- 适合关注生物安全的蛋白设计研究人员,尤其需规避意外毒性场景
蛋白质语言模型(PLMs)正成为从头设计蛋白的实用工具,但其双重用途带来安全风险。我们发现,对特定分类群进行领域适应会诱发有毒蛋白生成,即使毒性不是训练目标。为此,我们将Logit Diff Amplification(LDA)改造为PLMs的推理时控制机制。LDA通过放大基准模型与毒性微调模型之间的logit差值来调整词元概率,无需重新训练。在四个分类群上,LDA持续将预测毒性率(以ToxDL2衡量)降至分类群微调基线以下,同时保持生物学合理性。我们使用Fréchet ESM距离和预测折叠性(pLDDT)评估质量,结果表明LDA保持了与天然蛋白的分布相似性和结构可行性(优于依赖激活的引导方法,后者常导致序列质量下降)。结果证明,LDA为蛋白生成器提供了一种实用的安全调控手段,在降低诱导毒性的同时保留生成质量。
原文摘要 · Abstract (English)
Protein language models (PLMs) are becoming practical tools for de novo protein design, yet their dual-use potential raises safety concerns. We show that domain adaptation to specific taxonomic groups can elicit toxic protein generation, even when toxicity is not the training objective. To address this, we adapt Logit Diff Amplification (LDA) as an inference-time control mechanism for PLMs. LDA modifies token probabilities by amplifying the logit difference between a baseline model and a toxicity-finetuned model, requiring no retraining. Across four taxonomic groups, LDA consistently reduces predicted toxicity rate (measured via ToxDL2) below the taxon-finetuned baseline while preserving biological plausibility. We evaluate quality using Fréchet ESM Distance and predicted foldability (pLDDT), finding that LDA maintains distributional similarity to natural proteins and structural viability (unlike activation-based steering methods that tend to degrade sequence properties). Our results demonstrate that LDA provides a practical safety knob for protein generators that mitigates elicited toxicity while retaining generative quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。