让大模型一次生成多种人类价值观视角,提升回应多样性。
Overton Pluralistic Reinforcement Learning for Large Language Models
- 用双奖励机制训练模型,自动捕捉不同价值立场。
- 3B小模型在自然语言推理任务上比20B大模型高37.4%准确率。
- 无需额外模块或提示,适合需要多元观点的对话系统。
现有对齐范式难以捕捉人类价值观的多元性。过顿多元主义通过单次查询生成多重视角回应来弥补这一空白。本文提出OP-GRPO(过顿多元群体相对策略优化)强化学习框架,实现隐式过顿多元主义,使单一大语言模型在无显式提示或模块编排下生成多元回应。流程包含两步:首先,用句向量模型微调相似度估计器,更精准评估生成回应的覆盖范围;其次,将该估计器融入双奖励系统,确保真实人类视角广泛覆盖且各视角互异。实验证明“小模型、大视野”效应:训练后的Qwen2.5-3B-Instruct模型在自然语言推理基准上相比20B的GPT-OSS基线,相对准确率提升37.4%,同时优于模块化架构基线19.1%。进一步以GPT-4.1为判官的评估也证实方法稳健。
原文摘要 · Abstract (English)
Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism addresses this gap by generating responses with diverse perspectives from a single query. This paper introduces OP-GRPO (Overton Pluralistic Group Relative Policy Optimization), a reinforcement learning framework for implicit Overton Pluralism that enables a single large language model to produce pluralistic responses without explicit prompting or modular orchestration. Our workflow consists of two main steps. First, similarity estimator training fine-tunes a Sentence Transformer for Overton Pluralism tasks to provide more accurate coverage evaluation of generated responses. Second, OP-GRPO training incorporates this similarity estimator into a dual-reward system designed to ensure both broad coverage of genuine human perspectives and the uniqueness of each perspective, thereby promoting diversity. Empirical results demonstrate a "small models, big perspective coverage" effect. The trained Qwen2.5-3B-Instruct model surpasses a 20B GPT-OSS baseline with a 37.4 percent relative accuracy gain on a Natural Language Inference benchmark, and also outperforms a modular architecture baseline with a 19.1 percent relative improvement. Additional evaluations using GPT-4.1 as a large language model judge further confirm the robustness of the approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。