用因子图优化蛋白序列比对采样,可控保留进化信息。
Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

- 将序列采样建模为可控制的优化问题,用因子图推理选择序列。
- 在长程接触与构象预测任务中优于现有方法,提升结构敏感任务性能。
- 支持调节进化信号保留程度,适合需要可控进化信息的研究场景。
多序列比对(MSA)为蛋白质语言模型提供了明确的进化背景,但其深度大,在有限的词元预算下必须进行子采样。现有策略如随机选择、身份过滤和多样性驱动采样虽有效,但对保留的进化信号控制能力有限。本文将MSA子采样重述为显式优化问题,将查询身份与多样性等关键进化度量作为可控目标。基于此,提出AP-REASONER,一种基于亲和传播的因子图方法。通过进化感知的一元因子、代表一致性因子及两个控制参数,利用消息传递进行因子图推理,以推断固定预算下的MSA子集。在长程接触预测和构象集合预测任务上的实验表明,AP-REASONER在结构敏感下游任务中优于基线采样器,并能实现对替代蛋白构象的可控恢复。结果凸显了将MSA子采样建模为可控优化问题的价值,因子图推理为启发式选择提供了有效替代方案。
原文摘要 · Abstract (English)
Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based filtering, and diversity-driven sampling, are effective heuristics, yet provide limited control over the evolutionary signals retained in the subset. In this work, we recast MSA subsampling as an explicit optimization problem, where key evolutionary measures, including query identity and diversity, are treated as controllable objectives. Building on this view, we introduce AP-REASONER, an Affinity-Propagation-based factor-graph approach. With evolution-aware unary factors, exemplar-consistency factors, and two control knobs, AP-REASONER performs factor-graph reasoning through message passing to infer a fixed-budget MSA subset. Experiments on long-range contact prediction and conformational ensemble prediction show that AP-REASONER outperforms baseline subsamplers on structure-sensitive downstream tasks and enables controllable recovery of alternative protein conformations. These results highlight the value of modeling MSA subsampling as a controllable optimization problem, where factor-graph reasoning offers an effective alternative to heuristic selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。