arXiv:2602.03729cs.LG2026-02被引 1

用新正则化方法提升玻尔兹曼生成器训练效率,大幅减少数据需求。

Efficient Training of Boltzmann Generators Using Off-Policy Log-Dispersion Regularization

  • 引入离策略对数方差正则化,利用目标能量标签优化能量曲面形状。
  • 在多个基准上实现性能提升,数据效率最高提高十倍。
  • 适用于模拟数据或无目标样本的变分训练,通用性强。

从非归一化概率密度采样是计算科学的核心挑战。玻尔兹曼生成器是能够独立从给定温度下物理系统的玻尔兹曼分布中采样的生成模型。然而,其实际成功依赖于高效的数据训练,因为模拟数据和目标能量评估均成本高昂。为此,我们提出离策略对数方差正则化(LDR),一种基于对数方差目标推广的新正则化框架。将LDR与标准数据驱动训练目标结合,在无需额外在线策略样本的情况下应用于离策略设置。LDR通过利用目标能量标签中的附加信息,作为能量景观的形状正则化器。该正则化框架具有广泛适用性,支持无偏或有偏的模拟数据集,以及无需目标样本的纯变分训练。在所有基准测试中,LDR均提升了最终性能和数据效率,样本效率最高提升一个数量级。

原文摘要 · Abstract (English)

Sampling from unnormalized probability densities is a central challenge in computational science. Boltzmann generators are generative models that enable independent sampling from the Boltzmann distribution of physical systems at a given temperature. However, their practical success depends on data-efficient training, as both simulation data and target energy evaluations are costly. To this end, we propose off-policy log-dispersion regularization (LDR), a novel regularization framework that builds on a generalization of the log-variance objective. We apply LDR in the off-policy setting in combination with standard data-based training objectives, without requiring additional on-policy samples. LDR acts as a shape regularizer of the energy landscape by leveraging additional information in the form of target energy labels. The proposed regularization framework is broadly applicable, supporting unbiased or biased simulation datasets as well as purely variational training without access to target samples. Across all benchmarks, LDR improves both final performance and data efficiency, with sample efficiency gains of up to one order of magnitude.

生成模型能量正则化数据效率离策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。