提出流模型偏好优化中数据流形漂移问题及解决方案
Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

- 用温度控制机制约束偏好优化,防止样本偏离原始数据流形
- 在玩具基准上严格得分达0.899,优于FlowDPO的0.629
- 适合追求生成质量与分布一致性的扩散模型研究者
偏好优化是生成模型对齐的标准方法,但将其扩展到连续时间动态仍具挑战。在流匹配中,基于奖励的更新会改变传输轨迹,且缺乏对预训练数据流形的固有约束,可能导致终态样本脱离预训练支持集。我们将其失败模式形式化为流形漂移。理论上证明,最优流匹配可恢复终态数据分布,而偏好更新只要求终态位移存在非零法向分量,就会使样本离开预训练流形。为此,我们提出ThermoDPO,一种温度可控的目标函数,将成对偏好优化锚定在优选样本上。该目标在不同温度下连接拒绝采样微调与FlowDPO,并控制基于逐点重构的流形距离代理。为缓解低温下信号减弱问题,进一步引入加权版本ThermoDPO-weighted。在主要玩具基准上,ThermoDPO-weighted取得0.899的严格得分,高于FlowDPO的0.629和FlowDPO+RFT的0.857。在SD3.5-M、CFG=4.5下,其OCR提升47.5%,四项指标平均提升16.0%。
原文摘要 · Abstract (English)
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers the terminal data distribution, whereas a preference update leaves the pretrained manifold whenever its induced terminal displacement has a nonzero normal component. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO and controls a pointwise reconstruction-based surrogate for manifold distance. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted. On the main toy benchmark, ThermoDPO-weighted attains a StrictScore of 0.899, compared with 0.629 for FlowDPO and 0.857 for FlowDPO+RFT. On SD3.5-M at CFG = 4.5, it improves OCR by 47.5% and the average of four metrics by 16.0%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。