arXiv:2606.00609cs.LGcs.AI2026-06

解决大模型多领域强化学习中的奖励不可靠和能力冲突问题

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts

论文配图:CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
图 1 · 摘自论文原文
  • 用提示级评估协议生成可比奖励,应对非可验证任务
  • 通过方向感知能力投影,抑制跨领域能力冲突,提升优化效率
  • 在数学、对话等任务上显著优于基线,适配多领域模型训练

基于可验证奖励的强化学习在面向推理的大语言模型中取得显著进展,但将其扩展到多领域场景仍面临非可验证任务中奖励不可靠以及跨领域能力干扰的问题。本文提出CARE-RL,结合协议感知奖励生成与能力感知优化,缓解跨域冲突。针对非可验证任务,提出协议感知生成奖励模型(PA-GRM),在生成轨迹条件奖励前构建提示级评估协议与模板,实现开放响应的任务自适应且可比较的评估。针对多领域优化,设计方向感知能力子空间投影(DACSP),从历史强化学习阶段提取能力方向,通过增强对齐分量、抑制冲突分量并保留正交更新,实现更稳定的跨域优化。在数学、对话和指令遵循等多个基准上的实验表明,CARE-RL持续优于标准多领域强化学习基线,在Qwen2.5-7B和Qwen3-4B上分别达到47.9和50.7的总平均得分。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with verifiable rewards has achieved strong progress in reasoning-oriented LLMs, but extending it to multi-domain RL remains challenging due to reward unreliability in non-verifiable tasks and capability interference across domains. We propose CARE-RL to combine protocol-aware reward generation with capability-aware optimization for mitigating cross-domain conflicts. For non-verifiable tasks, the Protocol-Aware Generative Reward Model (PA-GRM) constructs prompt-level evaluation protocols and schemas before producing trace-conditioned rewards, enabling task-adaptive yet comparable evaluation of open-ended responses. For multi-domain optimization, Direction-Aware Capability Subspace Projection (DACSP) extracts historical capability directions from previous RL stages and modulates later updates by amplifying aligned components, suppressing conflicting components, and preserving orthogonal updates. Experiments across math, chat, and instruction-following benchmarks show that CARE-RL consistently outperforms standard multi-domain RL baselines, achieving Total Avg scores of 47.9 and 50.7 on Qwen2.5-7B and Qwen3-4B, respectively.

强化学习大模型多领域奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。