用提示引导反思,让大模型对齐更准更快
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- 引入外部模型生成反思提示,增强偏好信号对比度
- 训练样本减少40%仍达顶尖对齐效果,幻觉率显著下降
- 适合追求高效、稳定对齐的多模态模型研究者
直接偏好优化(DPO)已成为对齐大语言和视觉语言模型的有效轻量级方法。然而,标准DPO中选择与拒绝响应由同一策略生成,二者常共享相似错误且KL散度较小,导致学习信号弱、收敛慢且不稳定。为此,本文提出反射式偏好优化(RPO),在DPO框架中引入提示引导的反思机制。RPO利用外部模型识别幻觉来源并生成简洁的反思提示,构建具有更强对比性和清晰偏好信号的同策略偏好对。理论上,条件化提示可提升期望偏好间距,增加互信息,并保持样本效率,同时属于原策略分布族。实验表明,RPO以更少的训练样本和迭代次数实现优越对齐效果,显著降低幻觉率,在多模态基准测试中达到当前最优性能。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has emerged as a lightweight and effective alternative to Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning with AI Feedback (RLAIF) for aligning large language and vision-language models. However, the standard DPO formulation, in which both the chosen and rejected responses are generated by the same policy, suffers from a weak learning signal because the two responses often share similar errors and exhibit small Kullback-Leibler (KL) divergence. This leads to slow and unstable convergence. To address this limitation, we introduce Reflective Preference Optimization (RPO), a new framework that incorporates hint-guided reflection into the DPO paradigm. RPO uses external models to identify hallucination sources and generate concise reflective hints, enabling the construction of on-policy preference pairs with stronger contrastiveness and clearer preference signals. We theoretically show that conditioning on hints increases the expected preference margin through mutual information and improves sample efficiency while remaining within the policy distribution family. Empirically, RPO achieves superior alignment with fewer training samples and iterations, substantially reducing hallucination rates and delivering state-of-the-art performance across multimodal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。