arXiv:2410.08134cs.LGcs.AI2024-10ICLR被引 35

提出新方法让离散扩散模型按需生成,无需训练即可控制输出结果。

Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction

  • 通过贝叶斯后验采样实现对预训练模型的无仿真控制
  • 在图像、文本、蛋白质任务中均提升生成质量与多样性
  • 适用于不可导奖励函数,适合需要精准调控的应用场景

离散数据生成在从ChatGPT到蛋白质序列设计等应用中至关重要。但实际应用需对生成内容进行控制以满足特定属性或奖励指标,传统方法依赖强化学习人类反馈(RLHF)。本文研究了掩码扩散模型(MDMs)的可控生成问题,提出离散去噪后验预测(DDPP)框架,将控制任务转化为概率推断问题,通过学习从目标贝叶斯后验采样实现无仿真控制。该框架衍生出三种新目标,均无需仿真,可扩展至任意不可导奖励函数。实验验证其在像素级图像建模、基于文本奖励的模型对齐、以及蛋白质语言模型微调中的有效性;湿实验显示优化后的蛋白序列可实现瞬时表达。

原文摘要 · Abstract (English)

Generative modeling of discrete data underlies important applications spanning text-based agents like ChatGPT to the design of the very building blocks of life in protein sequences. However, application domains need to exert control over the generated data by steering the generative process - typically via RLHF - to satisfy a specified property, reward, or affinity metric. In this paper, we study the problem of steering Masked Diffusion Models (MDMs), a recent class of discrete diffusion models that offer a compelling alternative to traditional autoregressive models. We introduce Discrete Denoising Posterior Prediction (DDPP), a novel framework that casts the task of steering pre-trained MDMs as a problem of probabilistic inference by learning to sample from a target Bayesian posterior. Our DDPP framework leads to a family of three novel objectives that are all simulation-free, and thus scalable while applying to general non-differentiable reward functions. Empirically, we instantiate DDPP by steering MDMs to perform class-conditional pixel-level image modeling, RLHF-based alignment of MDMs using text-based rewards, and finetuning protein language models to generate more diverse secondary structures and shorter proteins. We substantiate our designs via wet-lab validation, where we observe transient expression of reward-optimized protein sequences.

扩散模型生成控制蛋白质设计非可导奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。