arXiv:2602.05993cs.LGcs.AI2026-02被引 19

提出Diamond Maps,让生成模型能快速适配用户偏好。

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

  • 用随机流映射设计,单步采样即可高效对齐奖励函数。
  • 在多个数据集上表现优于现有方法,且可扩展性更强。
  • 适合需要快速响应用户需求的生成应用开发者。

流模型与扩散模型虽能生成高质量样本,但训练后适应用户偏好或约束成本高且脆弱,此问题称为奖励对齐。我们认为高效的奖励对齐应是生成模型本身的属性,而非事后补救。为此提出“Diamond Maps”——一种随机流映射模型,可在推理时高效准确地对齐任意奖励函数。该模型将多步模拟过程压缩为单步采样,类似流映射,同时保留最优对齐所需的随机性。这一设计使搜索、序列蒙特卡洛和引导机制实现可扩展,支持高效一致的价值函数估计。实验表明,Diamond Maps 可通过从 GLASS Flows 中蒸馏学习,实现更强的奖励对齐性能,并优于现有方法,具有更好可扩展性。结果表明,这为生成模型在推理时快速适应任意偏好与约束提供了可行路径。

原文摘要 · Abstract (English)

Flow and diffusion models produce high-quality samples, but adapting them to user preferences or constraints post-training remains costly and brittle, a challenge commonly called reward alignment. We argue that efficient reward alignment should be a property of the generative model itself, not an afterthought, and redesign the model for adaptability. We propose "Diamond Maps", stochastic flow map models that enable efficient and accurate alignment to arbitrary rewards at inference time. Diamond Maps amortize many simulation steps into a single-step sampler, like flow maps, while preserving the stochasticity required for optimal reward alignment. This design makes search, Sequential Monte Carlo, and guidance scalable by enabling efficient and consistent estimation of the value function. Our experiments show that Diamond Maps can be learned efficiently via distillation from GLASS Flows, achieve stronger reward alignment performance, and scale better than existing methods. Our results point toward a practical route to generative models that can be rapidly adapted to arbitrary preferences and constraints at inference time.

生成模型奖励对齐流模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。