arXiv:2608.25855cs.CEcs.AI2026-08中稿 · EMNLP

改进蛋白质模型推理策略,无需重训练就能显著提升生成效果。

Unlocking Multimodal Protein Language Models at Inference Time

论文配图:Unlocking Multimodal Protein Language Models at Inference Time
图 1 · 摘自论文原文
  • 设计三阶段框架,系统测试不同推理采样方法
  • 在四项任务中实现显著性能提升,最高突破现有上限
  • 发现任务特异性采样偏好,挑战原有模型认知

多模态蛋白质语言模型(pLMs)学习蛋白序列与结构的联合分布,其生成性能也高度依赖推理时的采样策略。然而,先前研究更多关注模型训练,而忽视了推理策略的影响。本文建立一个三阶段实证研究框架,跨三种代表性pLMs和四项基础任务,评估原始采样、任务特定无分类器引导以及奖励引导束搜索等策略,分别控制采样分布、每步逻辑值和并行轨迹。通过探索-利用权衡的互补优化,我们(1)揭示默认推理协议的次优性,并识别出任务导向的采样偏好;(2)在各项任务中观察到显著的定量提升,无需更新模型参数即可持续提升多模态pLMs的上限性能;(3)得出关于基础模型的新结论,与以往共识相异。

原文摘要 · Abstract (English)

Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.

蛋白质生成推理优化多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。