让蛋白相互作用预测可解释,通过思维树和嵌入流匹配提升准确率与可信度
Protein Thoughts: Interpretable Reasoning with Tree of Thoughts and Embedding-Space Flow Matching for Protein-Protein Interaction Discovery

- 将蛋白互作拆解为四种生物信号,保留各因素独立贡献以实现透明推理
- 在SHS148k数据集上最优绑定者排名提升至11.2(基线47.7),性能提高76%
- 适合需要可解释性与机制验证的生物学家及药物研发人员
蛋白质-蛋白质相互作用(PPI)几乎调控所有细胞过程,但现有计算方法仅输出排序结果,缺乏机制解释,导致生物学家难以判断预测是否反映真实生化原理。本文提出「Protein Thoughts」框架,将PPI发现重构为可解释的搜索问题。系统将结合证据分解为四类生物意义信号:序列相似性(进化关系)、结构互补性(几何匹配)、界面平衡与化学兼容性(残基级互作)。不将其合并为黑箱得分,而是通过透明价值函数保留各信号贡献,支持排序与审计。为高效探索大规模候选空间,引入假设引导的熵正则化思维树搜索:微调语言模型基于嵌入特征生成高优先、探索或跳过指令,构建玻尔兹曼策略以平衡利用与熵驱动探索,同时假设感知剪枝避免过早放弃潜在优质候选。对得分不一致的候选,采用假设条件嵌入空间流匹配,将蛋白嵌入向结合物流形迁移。在SHS148k基准测试中,Protein Thoughts实现最优绑定者平均排名11.2(基线47.7),提升76%;绑定预测的值函数取得91.08±0.19的微平均F1,优于现有方法。
原文摘要 · Abstract (English)
Protein-protein interactions (PPIs) govern nearly all cellular processes, yet computational methods for identifying binding partners typically produce ranked predictions without mechanistic justification. This creates a fundamental barrier to adoption because biologists cannot assess whether predictions reflect genuine biochemical insight or spurious correlations. We present \textbf{Protein Thoughts}, a framework that reformulates PPI discovery as an interpretable search problem with explicit reasoning. The system decomposes binding evidence into four biologically meaningful signals: sequence similarity reflecting evolutionary relationships, structural complementarity capturing geometric fit, interface balance, and chemical compatibility encoding residue-level interactions. Rather than collapsing these signals into an opaque score, we preserve their individual contributions through a transparent value function that enables both ranking and auditing. To navigate large candidate spaces efficiently, we introduce hypothesis-guided entropy-regularized Tree-of-Thoughts search. A fine-tuned language model generates search directives from embedding-derived features, classifying candidates as high-priority, exploratory, or skippable. These directives condition a Boltzmann policy that balances exploitation with entropy-driven exploration, while hypothesis-aware pruning prevents premature abandonment of promising candidates. For candidates exhibiting score disagreement, hypothesis-conditioned embedding-space flow matching transports protein embeddings toward the binder manifold. On the SHS148k benchmark, Protein Thoughts achieves mean best-binder rank of 11.2 versus 47.7 for an entropic tree search baseline, a 76% improvement, and for binding prediction the trained value function achieves $91.08 \pm 0.19$ Micro-F1, outperforming existing PPI methods on the same dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。