用先验知识指导生成模型优化,提升视觉生成的准确性与稳定性。
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
- 引入贝叶斯先验锚点建模奖励不确定性,动态调节信任度。
- 在图像和视频生成任务中,语义对齐更强、收敛更快、感知质量更高。
- 适合追求高精度生成结果的研究者和工业应用开发者。
组相对策略优化(GRPO)已成为后训练视觉生成模型的有效且轻量级框架。然而,其性能受限于文本-视觉对应关系的模糊性:一个提示可能合理描述多种视觉输出,一张图像或视频也可能存在多个同等正确的解释。这种多对多关系导致奖励模型产生不确定且区分度弱的信号,使GRPO未能充分利用可靠反馈,反而过拟合噪声信号。本文提出贝叶斯先验引导优化(BPGO),作为GRPO的新扩展,通过语义先验锚点显式建模奖励不确定性。BPGO在两个层面自适应调节优化信任:组间贝叶斯信任分配强调与先验一致的组更新,抑制模糊组;组内先验锚定重归一化通过扩大可信偏差、压缩不确定得分,增强样本区分度。在图像与视频生成任务中,BPGO在语义对齐、感知保真度和收敛速度上均优于标准GRPO及近期变体。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual visual correspondence: a single prompt may validly describe diverse visual outputs, and a single image or video may support multiple equally correct interpretations. This many to many relationship leads reward models to generate uncertain and weakly discriminative signals, causing GRPO to underutilize reliable feedback and overfit noisy ones. We introduce Bayesian Prior-Guided Optimization (BPGO), a novel extension of GRPO that explicitly models reward uncertainty through a semantic prior anchor. BPGO adaptively modulates optimization trust at two levels: inter-group Bayesian trust allocation emphasizes updates from groups consistent with the prior while down-weighting ambiguous ones, and intra-group prior-anchored renormalization sharpens sample distinctions by expanding confident deviations and compressing uncertain scores. Across both image and video generation tasks, BPGO delivers consistently stronger semantic alignment, enhanced perceptual fidelity, and faster convergence than standard GRPO and recent variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。