arXiv:2604.17928cs.LGcs.AI2026-04ACL被引 1

用混合领域熵对齐,解决少样本强化学习中探索不足问题。

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment

论文配图:HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
图 1 · 摘自论文原文
  • 引入通用领域数据与熵动态对齐机制,提升探索多样性。
  • 仅用32个目标数据,性能媲美1000个样本的全量训练。
  • 适合资源受限场景下的大模型推理训练,尤其少样本任务。

基于可验证奖励的强化学习(RLVR)在训练面向推理的大语言模型方面表现优异,但现有方法多依赖高资源环境。在低资源场景下,RLVR易出现更严重的熵坍缩,严重限制探索并降低推理能力。为此,本文提出专为少样本RLVR设计的混合领域熵动态对齐(HEAL)框架。HEAL首先选择性引入高价值通用领域数据以促进多样化探索;随后提出熵动态对齐(EDA)奖励机制,对齐目标域与通用域在轨迹层面的熵动态,同时捕捉熵的幅度与细粒度变化。通过该对齐,EDA不仅缓解熵坍缩,还促使策略从通用域学习更多样化的行为。跨多个领域的实验表明,HEAL持续提升少样本RLVR性能。值得注意的是,仅使用32个目标域样本,其性能即可匹配甚至超越使用1000个目标样本的全量训练。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Reward (RLVR) has proven effective for training reasoning-oriented large language models, but existing methods largely assume high-resource settings with abundant training data. In low-resource scenarios, RLVR is prone to more severe entropy collapse, which substantially limits exploration and degrades reasoning performance. To address this issue, we propose Hybrid-domain Entropy dynamics ALignment (HEAL), a framework tailored for few-shot RLVR. HEAL first selectively incorporates high-value general-domain data to promote more diverse exploration. Then, we introduce Entropy Dynamics Alignment (EDA), a reward mechanism that aligns trajectory-level entropy dynamics between the target and general domains, capturing both entropy magnitude and fine-grained variation. Through this alignment, EDA not only further mitigates entropy collapse but also encourages the policy to acquire more diverse exploration behaviors from the general domain. Experiments across multiple domains show that HEAL consistently improves few-shot RLVR performance. Notably, using only 32 target-domain samples, HEAL matches or even surpasses full-shot RLVR trained with 1K target-domain samples.

强化学习少样本学习大模型推理熵对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。