arXiv:2605.27846cs.AI2026-05

通过自适应权重提升开放问答中策略优化的多样性和稳定性

EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA

  • 根据策略熵动态调整正负样本权重,避免熵崩溃
  • 在两个医疗开放问答数据集上显著提升响应多样性与稳定性
  • 适合需要高质量、多样化回答的开放问答场景

大型推理模型通常通过可验证奖励的强化学习(RLVR)进行训练。然而,现有方法对正负样本采用固定权重,结论难以推广到开放问答任务。本文系统研究了正负样本在开放问答强化学习中的作用:提出基于奖励均值的正负样本区分策略,发现负样本主导响应多样性与性能上限,正样本则决定响应质量与收敛稳定性。据此提出EAPO方法,基于当前策略熵与初始熵的比值自适应计算正样本权重:熵下降阶段降低正样本权重以保持探索,熵上升阶段增强权重以强化稳定,从而缓解熵崩溃问题。在两个公开的医疗开放问答数据集上的实验表明,EAPO在响应多样性与稳定性方面持续显著优于固定权重基线。

原文摘要 · Abstract (English)

Large Reasoning Models are typically trained via reinforcement learning from verifiable rewards (RLVR). However, existing approaches adopt fixed weights for positive and negative samples, and the conclusions hardly generalize to open-ended question answering (QA). In this paper, we systematically investigate the roles of positive and negative samples in reinforcement learning for open-ended QA. We propose a reward-mean-based strategy for distinguishing positive from negative samples, and observe that negative samples predominantly govern response diversity and the performance upper bound, whereas positive samples primarily determine response quality and convergence stability. Building on these observations, we propose EAPO, an Entropy-driven Adaptive Policy Optimization method that adaptively computes the weighting coefficients of positive samples based on the ratio of the current policy entropy to the initial entropy. During the entropy-decreasing phase, the weight assigned to positive samples is reduced to preserve exploration, whereas during the entropy-increasing phase it is amplified to reinforce stability, thereby mitigating entropy collapse. Experiments on two publicly available open-ended medical QA datasets demonstrate that EAPO consistently and substantially outperforms fixed-weight baselines in both response diversity and stability.

强化学习开放问答策略优化熵控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。