arXiv:2603.10588cs.AIcs.CL2026-03被引 1

实验证明道德推理无需刻意追求多样性,标准强化学习更有效。

Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning

  • 用评分模型构建稳定奖励管道,对比分布匹配与奖励最大化方法
  • 道德推理中高分答案聚集度高,多样解法得分差异小
  • 适合关注大模型对齐、强化学习应用的研究者

强化学习结合可验证奖励(RLVR)在逻辑推理任务中表现优异,但大语言模型对齐是否需要根本不同的方法仍不明确。鉴于道德推理中多种合理回答并存,自然假设是对齐任务需依赖多样性导向的分布匹配算法,而非奖励最大化的策略方法。本研究首次在MoReBench数据集上系统比较两种范式。为实现稳定RLVR训练,我们基于评分标准训练了一个Qwen3-1.7B判别模型构成奖励管道。出乎意料的是,分布匹配方法在对齐任务上并未显著优于奖励最大化方法。通过将高奖励响应映射至语义空间进行可视化分析,发现道德推理的高分分布比数学推理更集中,不同解法策略往往获得相似高分。这一反直觉结果解释了为何模式搜索优化在对齐任务中同样甚至更有效。结果表明,对齐任务并非本质上需要保留多样性的算法,标准奖励最大化型RLVR方法可无需显式多样性机制直接迁移至道德推理。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamentally different approaches remains unclear. Given the apparent tolerance for multiple valid responses in moral reasoning, a natural hypothesis is that alignment tasks inherently require diversity-seeking distribution-matching algorithms rather than reward-maximizing policy-based methods. We conduct the first comprehensive empirical study comparing both paradigms on MoReBench. To enable stable RLVR training, we build a rubric-grounded reward pipeline by training a Qwen3-1.7B judge model. Contrary to our hypothesis, we find that distribution-matching approaches do not demonstrate significant advantages over reward-maximizing methods as expected on alignment tasks. Through semantic visualization mapping high-reward responses to semantic space, we demonstrate that moral reasoning exhibits more concentrated high-reward distributions than mathematical reasoning, where diverse solution strategies yield similarly high rewards. This counter-intuitive finding explains why mode-seeking optimization proves equally or more effective for alignment tasks. Our results suggest that alignment tasks do not inherently require diversity-preserving algorithms, and standard reward-maximizing RLVR methods can effectively transfer to moral reasoning without explicit diversity mechanisms.

大模型对齐强化学习道德推理奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。