arXiv:2603.20115cs.LGq-bio.BM2026-03被引 3

通过调整记忆多重性,实现蛋白质生成的精准定向。

Conditioning Protein Generation via Hopfield Pattern Multiplicity

  • 用多重性比率直接控制生成方向,无需训练。
  • 在5个Pfam家族中准确匹配目标序列分布,关键残基恢复依赖PCA效果。
  • 适合需要快速生成特定蛋白子集的研究者,如药物设计候选探索。

小蛋白家族比对中常包含一个感兴趣子集,但标注数据不足难以训练条件生成器。本文通过向随机注意力采样器的logits添加一个多重性比率来实现无训练条件化,增大该比率可使生成结果从全家族转向指定子集。当记忆归一化时,对应的玻尔兹曼分布恰好为高斯混合模型,各成分权重由多重性决定。该方法将潜在空间中的精确条件化与采样、PCA重构和序列解码带来的误差分离。在五个Pfam家族中,注意力机制符合解析目标;但单残基标记恢复效果取决于PCA对目标与背景序列的分离能力。匹配加权轮廓HMM能更直接重现这些标记,而随机注意力在Kunitz比较中展现出更低的ESM2伪困惑度。以23条经筛选的omega-毒素序列为目标子集,生成了保留半胱氨酸骨架和Tyr13的多样化序列,其余残基向目标集偏移。这些序列可作为实验测试候选,但未证明其结合能力。

原文摘要 · Abstract (English)

Small protein-family alignments often contain a subset of interest but not enough labeled data to train a conditional generator. We condition a training-free stochastic-attention sampler by adding one multiplicity ratio to its logits. Increasing this ratio shifts generation from the full family toward the designated subset. For unit-norm memories, the resulting Boltzmann distribution is exactly a Gaussian mixture whose component weights are set by the multiplicities. This result separates exact conditioning in latent space from losses caused by sampling, PCA reconstruction, and sequence decoding. Across five Pfam families, attention followed the analytic target, but recovery of single-residue markers depended on how well PCA separated the designated and background sequences. A matched weighted profile HMM reproduced these markers more directly, while stochastic attention gave lower ESM2 pseudo-perplexity in the Kunitz comparison. Using a curated set of 23 omega-conotoxin sequences as the target subset produced diverse sequences that preserved the cysteine scaffold and Tyr13 and shifted other residues toward the designated set. These sequences are candidates for experimental testing; they do not establish binding.

蛋白质生成条件采样无训练生成多重性调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。