arXiv:2509.17937cond-mat.softcs.LG2025-09

用随机非线性投影压缩分子数据,提速分析且不丢关键信息。

Random functions as data compressors for machine learning of molecular processes

  • 用随机非线性映射压缩高维分子特征空间。
  • 在NTL9和双诺亮氨酸病毒头片测试中保持静态与动态特性。
  • 适合需要加速轨迹分析的蛋白质折叠研究者。

机器学习正快速改变分子动力学模拟的执行与分析方式,涵盖材料建模到蛋白质折叠与功能研究。许多算法需选取少量描述问题的关键特征。尽管深度神经网络可处理大量输入特征,但训练成本随输入规模上升,因此对多数实际问题,特征子集选择至关重要。本文表明,随机非线性投影可用于压缩大特征空间,加快计算速度且几乎不损失信息。我们提出一种高效生成随机投影的方法,并以蛋白质折叠为例展示通用流程。在NTL9及双诺亮氨酸型病毒头片的测试中,随机压缩保留了原始高维特征空间的核心静态与动态信息,使轨迹分析更具鲁棒性。

原文摘要 · Abstract (English)

Machine learning (ML) is rapidly transforming the way molecular dynamics simulations are performed and analyzed, from materials modeling to studies of protein folding and function. ML algorithms are often employed to learn low-dimensional representations of conformational landscapes and to cluster trajectories into relevant metastable states. Most of these algorithms require selecting a small number of features that describe the problem of interest. Although deep neural networks can tackle large numbers of input features, the training costs increase with input size, which makes the selection of a subset of features mandatory for most problems of practical interest. Here, we show that random nonlinear projections can be used to compress large feature spaces and make computations faster without substantial loss of information. We describe an efficient way to produce random projections and then exemplify the general procedure for protein folding. For our test cases NTL9 and the double-norleucin variant of the villin headpiece, we find that random compression retains the core static and dynamic information of the original high dimensional feature space and makes trajectory analysis more robust.

分子模拟特征压缩蛋白质折叠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。