用热力学原理指导生成模型,在数据少时高效采样分子构象变化。
Minimum-Excess-Work Guidance
- 基于热力学功最小化设计正则化框架,结合最优传输思想。
- 在有限实验数据下显著减少偏差,提升稀疏数据场景的采样效率。
- 适合分子模拟、生物物理等数据稀缺领域的生成建模任务。
我们提出一种受热力学功启发的正则化框架,用于引导预训练的概率流生成模型(如连续归一化流或扩散模型),通过最小化过量功来实现高效引导。该方法适用于科学应用中常见的稀疏数据场景,即仅有少量目标样本或部分密度约束。提出两种策略:路径引导用于集中概率质量以采样罕见过渡态;可观测引导用于对齐生成分布与实验可观测量,同时保持熵。在粗粒度蛋白质模型上验证,成功引导模型采样折叠/展开状态间的过渡构象,并利用实验数据修正系统性偏差。该方法将热力学原理与现代生成架构结合,为数据稀缺领域提供了有理论基础、高效且物理启发的替代方案。实证结果表明样本效率提升且偏差降低,凸显其在分子模拟等领域的适用性。
原文摘要 · Abstract (English)
We propose a regularization framework inspired by thermodynamic work for guiding pre-trained probability flow generative models (e.g., continuous normalizing flows or diffusion models) by minimizing excess work, a concept rooted in statistical mechanics and with strong conceptual connections to optimal transport. Our approach enables efficient guidance in sparse-data regimes common to scientific applications, where only limited target samples or partial density constraints are available. We introduce two strategies: Path Guidance for sampling rare transition states by concentrating probability mass on user-defined subsets, and Observable Guidance for aligning generated distributions with experimental observables while preserving entropy. We demonstrate the framework's versatility on a coarse-grained protein model, guiding it to sample transition configurations between folded/unfolded states and correct systematic biases using experimental data. The method bridges thermodynamic principles with modern generative architectures, offering a principled, efficient, and physics-inspired alternative to standard fine-tuning in data-scarce domains. Empirical results highlight improved sample efficiency and bias reduction, underscoring its applicability to molecular simulations and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。