arXiv:2608.19808cs.LG2026-08

提升环肽设计可行性与生成鲁棒性,突破现有方法低成功率瓶颈

FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design

论文配图:FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design
图 1 · 摘自论文原文
  • 基于可行性感知的偏好构建与难度自适应优化
  • 在固定生成预算下成功率提升至57.79%(原46.89%)
  • 适合需要高可行性和多目标平衡的复杂靶点环肽设计

环肽因其高结合亲和力和结构稳定性,正成为药物发现中的有前景分子骨架。然而,将生成模型从线性肽扩展到环肽设计仍具挑战,因环化带来几何与生物物理约束的强耦合,大幅压缩可行设计空间。现有方法受限于训练数据不足,依赖零样本生成或后处理过滤,导致可行设计产出率低且难以控制多目标权衡。为此,我们提出FAR-DPO(可行性感知且鲁棒的直接偏好优化),一种与架构无关的框架,引导生成模型产出结构与生物物理上可行的环肽设计,尤其适用于困难靶点。FAR-DPO融合可行性感知偏好构建与难度感知组鲁棒优化:通过可行性门控多目标占优构建同靶点偏好对,并根据当前偏好损失自适应重加权预定义难度组。在CPSea LNR基准上,固定生成预算下,FAR-DPO将PepGLAD的成功率从46.89%提升至57.79%,PepFlow从47.96%提升至49.57%。增益同样体现在最困难靶点四分位,且伴随更优的最优靶点结合得分。结果表明FAR-DPO在提升可行性与靶点鲁棒性方面具有显著效果。

原文摘要 · Abstract (English)

Cyclic peptides are emerging as promising molecular scaffolds in drug discovery due to their high binding affinity and structural stability. However, extending generative models from linear to cyclic peptide design remains challenging, as cyclization sharply restricts the feasible design space through coupled geometric and biophysical constraints. Moreover, limited training data has led existing approaches to rely largely on zero-shot generation or post hoc filtering, resulting in low yields of feasible designs and limited control over multi-objective trade-offs. To address these limitations, we propose FAR-DPO (Feasibility-Aware and Robust Direct Preference Optimization), an architecture-agnostic framework that steers generative models toward structurally and biophysically feasible cyclic peptide designs, particularly for challenging targets. FAR-DPO integrates feasibility-aware preference construction with difficulty-aware group-robust optimization. Specifically, it constructs within-target preference pairs through feasibility-gated multi-objective dominance and adaptively reweights predefined difficulty groups according to their current preference losses. On the CPSea LNR benchmark, under a fixed generation budget, FAR-DPO increases overall success rate from 46.89% to 57.79% on PepGLAD and from 47.96% to 49.57% on PepFlow. These gains also extend to the hardest target quartile and are accompanied by more favorable best-per-target binding scores. Together, these results demonstrate FAR-DPO's effectiveness in improving feasibility and target-wise robustness.

环肽设计生成模型偏好优化药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。