打通序列功能预测的函数空间与权重空间,统一解释正则化与先验的关系。
On learning functions over biological sequence space: relating Gaussian process priors, regularization, and gauge fixing
- 用高斯过程与正则化回归建立序列功能映射的统一框架
- 揭示权重正则化如何隐含定义函数空间先验与唯一表示(规约)
- 可高效计算序列贡献度等统计量的后验分布,适合生物序列建模研究者
从生物序列(如DNA、RNA、蛋白质)到功能量化指标的映射在现代生物学中至关重要。本文关注两个任务:(i) 推断序列-功能映射模型,(ii) 分解映射以揭示子序列的贡献。由于每个映射在多重权重表示下不唯一,需通过“规约”(gauge fixing)定义唯一表示。已有研究表明,大多数规约形式是过参数化权重空间中L2正则化回归的唯一解,正则化项即对应规约方式。本文建立权重空间正则化与函数空间(即有限序列集上的实值函数空间)高斯过程方法之间的联系。我们阐明:权重空间正则化不仅施加函数空间的隐式先验,还限制最优权重落在特定规约中。我们提出构造方法,使正则化对应任意显式高斯过程先验与多种规约。同时刻画了常见正则化对应的隐式函数空间先验。最后,推导出一大类序列-功能统计量(包括规约后的权重及高阶上位性系数表达式)的后验分布,并证明对于乘积核先验,可通过核技巧高效计算。
原文摘要 · Abstract (English)
Mappings from biological sequences (DNA, RNA, protein) to quantitative measures of sequence functionality play an important role in contemporary biology. We are interested in the related tasks of (i) inferring predictive sequence-to-function maps and (ii) decomposing sequence-function maps to elucidate the contributions of individual subsequences. Because each sequence-function map can be written as a weighted sum over subsequences in multiple ways, meaningfully interpreting these weights requires ``gauge-fixing,'' i.e., defining a unique representation for each map. Recent work has established that most existing gauge-fixed representations arise as the unique solutions to $L_2$-regularized regression in an overparameterized ``weight space'' where the choice of regularizer defines the gauge. Here, we establish the relationship between regularized regression in overparameterized weight space and Gaussian process approaches that operate in ``function space,'' i.e.~the space of all real-valued functions on a finite set of sequences. We disentangle how weight space regularizers both impose an implicit prior on the learned function and restrict the optimal weights to a particular gauge. We show how to construct regularizers that correspond to arbitrary explicit Gaussian process priors combined with a wide variety of gauges and characterize the implicit function space priors associated with the most common weight space regularizers. Finally, we derive the posterior distribution of a broad class of sequence-to-function statistics, including gauge-fixed weights and multiple systems for expressing higher-order epistatic coefficients. We show that such distributions can be efficiently computed for product-kernel priors using a kernel trick.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。