提出自适应优先级机制,解决多目标图像生成中的优化失衡问题。
APEX: Learning Adaptive Priorities for Multi-Objective Alignment in Vision-Language Generation
- 采用双阶段自适应归一化稳定异构奖励信号
- 动态调度目标,实现四个指标的均衡提升
- 适合需要多目标平衡的视觉语言生成任务
文本到图像生成中的多目标对齐通常通过静态线性加权实现,但固定权重在异构奖励下常导致优化失衡:模型过度拟合高方差、高响应度的目标(如OCR),而忽视感知目标。我们识别出两个机制原因:方差劫持,即奖励分散导致隐式重加权主导训练信号;梯度冲突,即竞争目标产生相反更新方向,引发类似摇摆的振荡。为此提出APEX(基于自适应优先级的高效多目标对齐),通过双阶段自适应归一化稳定异构奖励,并利用P^3自适应优先级动态调度目标,综合学习潜力、冲突惩罚和进展需求。在Stable Diffusion 3.5上,APEX实现了四个异构目标的改进帕累托前沿,平衡提升+1.31 PickScore、+0.35 DeQA、+0.53 Aesthetics,同时保持竞争力的OCR准确率,有效缓解多目标对齐的不稳定性。
原文摘要 · Abstract (English)
Multi-objective alignment for text-to-image generation is commonly implemented via static linear scalarization, but fixed weights often fail under heterogeneous rewards, leading to optimization imbalance where models overfit high-variance, high-responsiveness objectives (e.g., OCR) while under-optimizing perceptual goals. We identify two mechanistic causes: variance hijacking, where reward dispersion induces implicit reweighting that dominates the normalized training signal, and gradient conflicts, where competing objectives produce opposing update directions and trigger seesaw-like oscillations. We propose APEX (Adaptive Priority-based Efficient X-objective Alignment), which stabilizes heterogeneous rewards with Dual-Stage Adaptive Normalization and dynamically schedules objectives via P^3 Adaptive Priorities that combine learning potential, conflict penalty, and progress need. On Stable Diffusion 3.5, APEX achieves improved Pareto trade-offs across four heterogeneous objectives, with balanced gains of +1.31 PickScore, +0.35 DeQA, and +0.53 Aesthetics while maintaining competitive OCR accuracy, mitigating the instability of multi-objective alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。