用判别式奖励提升视觉模型的感知判断力
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
- 通过判别式奖励收集多样负样本,再进行列表偏好优化
- 显著提升多模态大模型的视觉辨别能力,保持生成优势
- 适合关注视觉感知对齐与奖励机制设计的研究者
本文提出感知偏好优化(PerPO),一种针对生成式预训练多模态大语言模型(MLLMs)视觉辨别挑战的对齐方法。为使MLLMs更贴近人类视觉感知过程,PerPO采用判别式奖励获取多样化负样本,并通过列表偏好优化进行排序。利用奖励作为量化排序边际,该方法有效衔接生成偏好优化与判别式经验风险最小化。PerPO显著增强MLLMs的视觉辨别能力,同时保留其生成优势,缓解图像无关奖励欺骗问题,并确保各类视觉任务中表现一致。此项工作推动了更具感知对齐性与通用性的MLLMs发展,也呼吁社区重新思考MLLM对齐策略。
原文摘要 · Abstract (English)
This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained multimodal large language models (MLLMs). To align MLLMs with human visual perception process, PerPO employs discriminative rewarding to gather diverse negative samples, followed by listwise preference optimization to rank them.By utilizing the reward as a quantitative margin for ranking, our method effectively bridges generative preference optimization and discriminative empirical risk minimization. PerPO significantly enhances MLLMs' visual discrimination capabilities while maintaining their generative strengths, mitigates image-unconditional reward hacking, and ensures consistent performance across visual tasks. This work marks a crucial step towards more perceptually aligned and versatile MLLMs. We also hope that PerPO will encourage the community to rethink MLLM alignment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。