提出新模型AgForce,让抗体生成更精准响应抗原靶点。
AgForce Enables Antigen-conditioned Generative Antibody Design

- 用图神经网络和特殊解码器联合设计抗体序列与结构。
- 在多个指标上超越基线,有效词汇量提升约2倍。
- 适合需要精准抗原响应的抗体药物研发人员。
现有抗体设计方法虽以抗原结构为条件生成互补决定区(CDR),但系统评估发现其大多忽略抗原输入。研究识别出三种失效模式:抗原盲视源于模型依赖抗体框架上下文而非抗原信息,导致不同靶点生成几乎相同的CDR;词汇坍塌使预测氨基酸减少至每位置3~5种,远低于天然序列分布;交叉熵损失的贪心解码存在argmax瓶颈,仅依赖位置与环长,无视抗原目标。为此提出新型编码器-解码器架构AgForce,采用图神经网络(GNN)作为编码器,结合框架丢弃、门控瓶颈与双曲交叉注意力抑制抗体捷径路径。解码器引入混合密度网络(MDN)序列头,包含类Potts配对耦合与退火多选学习(aMCL),替代交叉熵目标,生成多成分分布。抗原循环一致性头将梯度回传至序列解码器,强制预测分布编码抗原身份。在ChiMERa-Bench数据集上,AgForce在所有三组划分中,于fnat、iRMSD、DockQ及表位F1等接口指标上全面领先基线,同时实现约2倍于GNN基线的有效词汇量。
原文摘要 · Abstract (English)
Antibody design methods condition on antigen structure to generate complementarity-determining regions (CDR), yet a systematic evaluation of baseline methods reveals that they largely ignore the antigen input. We identify three failure modes that explain this behavior. Antigen blindness arises because models derive predictions from antibody framework context rather than antigen information, producing nearly identical CDRs regardless of the target. Vocabulary collapse reduces predicted amino acids to three to five per position, far below the ground truth distribution in native sequences. The argmax bottleneck on the cross-entropy loss via greedy decoding acts as a null predictor that sees only position and loop length and recovers almost all the residues irrespective of the model architecture and the target antigen. We propose a novel encoder-decoder architecture called AgForce that uses a graph neural network (GNN) as the encoder and specialized decoders for sequence-structure co-design. Specifically, we apply framework dropout, gated bottlenecks, and hyperbolic cross attention that prevent the antibody shortcut path. In the decoder, a Mixture Density Network (MDN) sequence head with Potts-like pairwise coupling and annealed Multiple Choice Learning (aMCL) replaces the cross-entropy objective with a multi-component distribution. An antigen cycle consistency head routes gradients through the sequence decoder, forcing predicted distributions to encode antigen identity. On the ChiMERa-Bench dataset, AgForce leads the baselines on every interface metric on all three splits (fnat, iRMSD, DockQ, and epitope F1) together with structure and sequence, while achieving roughly a 2x larger effective vocabulary compared to the GNN baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。