arXiv:2510.09784cs.LGcond-mat.stat-mech2025-10被引 1

将分子表示学习与生成结合,实现高效低维表征。

Combined Representation and Generation with Diffusive State Predictive Information Bottleneck

  • 用时滞信息瓶颈+扩散模型联合训练,统一表示与生成目标。
  • 可融合多条模拟轨迹温度信息,学习热力学一致表征。
  • 适合分子生成、物性预测等数据稀缺场景研究者使用。

生成建模在高维空间中越来越依赖大量数据。在分子科学中,数据采集成本高且关键事件稀少,因此将数据压缩到低维流形对下游任务(包括生成)尤为重要。本文提出一种联合框架——扩散状态预测信息瓶颈(D-SPIB),将时滞信息瓶颈用于捕捉分子重要表征,并与扩散模型在单一联合目标下联合训练。该方法使表示学习与生成目标在灵活架构中实现平衡。此外,模型能整合来自不同分子模拟轨迹的温度信息,学习一致且有用的热力学内部表征。我们在多个分子任务上对D-SPIB进行基准测试,展示其在训练集外物理条件下探索的潜力。

原文摘要 · Abstract (English)

Generative modeling becomes increasingly data-intensive in high-dimensional spaces. In molecular science, where data collection is expensive and important events are rare, compression to lower-dimensional manifolds is especially important for various downstream tasks, including generation. We combine a time-lagged information bottleneck designed to characterize molecular important representations and a diffusion model in one joint training objective. The resulting protocol, which we term Diffusive State Predictive Information Bottleneck (D-SPIB), enables the balancing of representation learning and generation aims in one flexible architecture. Additionally, the model is capable of combining temperature information from different molecular simulation trajectories to learn a coherent and useful internal representation of thermodynamics. We benchmark D-SPIB on multiple molecular tasks and showcase its potential for exploring physical conditions outside the training set.

分子生成扩散模型表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。