提升主观问题推理多样性,让模型从多角度给出更合理答案。
Diversity-Enhanced Reasoning for Subjective Questions
- 用多角色视角合成推理数据,增强答案多样性
- 在主观任务上准确率提升14.1%(域内)和7.64%(域外)
- 多样性比推理长度更能反映模型真实表现,适合需要多角度判断的任务
具备长链思维能力的大规模推理模型(LRMs)在数学求解和代码生成等客观任务中表现优异,但通过可验证奖励强化学习(RLVR)优化后,其生成多样性下降,难以应对具有多重答案的主观推理任务。本文发现,引入角色多样性与词级多样性可有效提升主观推理性能:前者以真实利益相关者为锚点提供连贯框架,后者拓展答案搜索空间。我们提出MultiRole-R1框架,包含无监督数据构建流程,合成融合多角色视角的推理链,并采用基于组相对策略优化的强化学习,将多样性作为奖励信号之一。仅在主观任务上训练,MultiRole-R1使域内和域外准确率分别提升14.1%和7.64%,甚至改善了AIME 2024等高级数学推理表现。结果还表明,多样性是比推理长度更稳定的准确率指标。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) with long chain-of-thought capabilities, optimized via reinforcement learning with verifiable rewards (RLVR), excel at objective reasoning tasks like mathematical problem solving and code generation. However, RLVR is known for degrading generation diversity, which causes LRMs to fall short on subjective reasoning that has multiple answers depending on different role perspectives. While recent studies recognize the importance of diversity-enhanced training in objective reasoning, limited attention has been given to subjective tasks. In this paper, we find that subjective reasoning can be improved by introducing perspective diversity and token-level diversity, with the former one providing a coherent scaffolding anchored to a real-world stakeholder group and the latter one broadening the answer search space. We propose MultiRole-R1, a diversity-enhanced training framework featuring an unsupervised data construction pipeline that synthesizes reasoning chains incorporating various role perspectives. It also employs reinforcement learning via Group Relative Policy Optimization with reward shaping, taking diversity as a reward signal in addition to verifiable reward. Training on subjective tasks solely, MultiRole-R1 increases the in-domain and out-of-domain accuracy by 14.1% and 7.64%, and even enhances the performance on advanced math reasoning such as AIME 2024. We further show that diversity is a more consistent indicator of accuracy than reasoning length.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。