提出可同时处理在线与离线数据的生成模型训练损失族,提升稳定性与泛化性。
$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

- 构建基于f-散度的损失家族,支持在线与离线数据统一训练
- 在分子生成与大语言模型调优中验证了性能稳定性和模式覆盖能力
- 适用于GFlowNets、变分推断及大模型调优,兼顾理论严谨性与实用价值
在GFlowNets和变分推断中,目标与模型对数概率间的均方误差被证明是一种有效且方差低的替代损失。该损失在在线策略下梯度等价于KL散度,在离线策略下仍保持有效并具有相同的全局最小值。本文展示该构造可推广至整个f-散度家族,形成一类损失函数:其在线策略梯度对应相应f-散度,而离线策略下仍保持相同全局最小值。我们证明,目标与模型对数概率上平移不变的损失函数与f-散度之间存在一一对应关系。这使得我们能设计出继承对应f-散度特性的新代理损失函数(如更优的模式覆盖),并适用于离线数据。我们在多种任务中验证该方法,包括经典合成示例、用于分子发现的SynFlowNets,以及异步大语言模型(LLM)调优,结果表明模型在在线与离线策略下均保持预测特性,适用范围涵盖广泛的生成模型。
原文摘要 · Abstract (English)
In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated \emph{on-policy} its gradients correspond to those of the KL divergence, while \emph{off-policy} it remains a valid loss with the same global minimizer. In this work, we demonstrate that this construction can be extended to the whole family of $f$-divergences, leading to a family of losses whose on-policy gradients are that of the corresponding $f$-divergence, but retain the same global minimizer off-policy. Specifically, we show that the on-policy gradients lead to a one to one correspondence between translation invariant loss functions on the target and model log probabilities, and $f$-divergences. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that inherit the properties of the corresponding $f$-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including classic synthetic examples, SynFlowNets for molecule discovery, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。