arXiv:2602.00782q-bio.BMcs.AI2026-02被引 5

解决蛋白质语言模型生成时的重复问题,提升结构可靠性。

Controlling Repetition in Protein Language Models

  • 提出量化指标评估基序和均聚物重复,揭示其对折叠可靠性的影响。
  • 设计UCCS方法,通过约束数据集对比学习,精准抑制重复而不损害折叠能力。
  • 无需重训练或启发式解码,适用于ESM-3、ProtGPT2等主流模型。

蛋白质语言模型(PLMs)在结构预测和从头蛋白设计中取得进展,但在生成过程中常出现病理级重复问题。与文本中仅影响可读性不同,蛋白质中的重复会削弱结构置信度和功能可行性。本文首次系统研究了PLMs中的重复现象,提出定量指标衡量基序级和均聚物重复,并证明其对折叠可靠性有显著负面影响。为此,提出UCCS(Utility-Controlled Contrastive Steering)方法,通过构建在结构功能上保持一致但重复程度差异大的对比数据集,实现重复与结构可用性的解耦。该方法生成的引导向量可在推理阶段注入,持续降低重复率,且无需重新训练或使用启发式解码。在CATH、UniRef50和SCOP数据集上的实验表明,UCCS优于解码惩罚和其他基线方法,显著减少重复同时维持AlphaFold置信度。结果确立重复控制是PLMs的核心挑战,并强调数据驱动的引导策略为可靠蛋白生成的可行路径。

原文摘要 · Abstract (English)

Protein language models (PLMs) have enabled advances in structure prediction and de novo protein design, yet they frequently collapse into pathological repetition during generation. Unlike in text, where repetition merely reduces readability, in proteins it undermines structural confidence and functional viability. To unify this problem, we present the first systematic study of repetition in PLMs. We first propose quantitative metrics to characterize motif-level and homopolymer repetition and then demonstrate their negative impact on folding reliability. To address this challenge, we propose UCCS (Utility-Controlled Contrastive Steering), which steers protein generation with a constrained dataset. Instead of naively contrasting high- vs. low-repetition sequences, we construct contrastive sets that maximize differences in repetition while tightly controlling for structural utility. This disentanglement yields steering vectors that specifically target repetition without degrading foldability. Injected at inference, these vectors consistently reduce repetition without retraining or heuristic decoding. Experiments with ESM-3 and ProtGPT2 in CATH, UniRef50, and SCOP show that our method outperforms decoding penalties and other baselines, substantially lowering repetition while preserving AlphaFold confidence scores. Our results establish repetition control as a central challenge for PLMs and highlight dataset-guided steering as a principled approach for reliable protein generation.

蛋白质生成语言模型重复控制结构预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。