通过动态采样提升规则推理模型在复杂场景下的表现与效率。
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
- 基于历史奖励动态调整领域权重,实现自适应任务采样。
- 在8个分布内任务上比OpenAI-o1高4.1%,3个分布外任务上高10.4%。
- 无需人工设计训练混合策略,适合多领域规则推理应用。
基于规则的推理被认为是推理的核心问题之一。尽管大型推理模型(LRMs)在强化学习(RL)增强下展现出卓越的推理能力,但实际应用仍面临规则格式、类型和复杂性差异带来的严峻挑战。为此,我们提出RuleReasoner,一种基于大量精心构建任务和新颖领域感知动态采样机制的规则推理方法。具体而言,RuleReasoner通过根据历史奖励更新领域权重,对每个训练批次进行重采样,实现领域平衡与主动学习调度,避免了依赖人工设计的静态混合训练。在分布内(ID)和分布外(OOD)基准上的评估表明,RuleReasoner显著优于前沿的LRMs:在八个ID任务上比OpenAI-o1高出4.1%,在三个OOD任务上高出10.4%。此外,该方法还表现出更高的计算效率。
原文摘要 · Abstract (English)
Rule-based reasoning is acknowledged as one of the fundamental problems of reasoning. While recent studies show that large reasoning models (LRMs) have remarkable reasoning capabilities enhanced by reinforcement learning (RL), real applications still face severe challenges due to variations in rule formats, types, and complexity. To mitigate this issue, we introduce RuleReasoner, an effective method for rule-based reasoning via a wide collection of curated tasks and a novel domain-aware dynamic sampling approach in RL. Specifically, RuleReasoner resamples each training batch by updating the domain weights based on historical rewards. This facilitates domain balance and active learning schedules for RL, obviating static mix-training engineered by human. Evaluations on in-distribution (ID) and out-of-distribution (OOD) benchmarks reveal that RuleReasoner outperforms frontier LRMs by a significant margin ($Δ$4.1% on eight ID tasks and $Δ$10.4% on three OOD tasks over OpenAI-o1). Notably, our approach also exhibits higher computational efficiency compared to prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。