arXiv:2506.13351cs.CLcs.AI2025-06被引 4

让大模型在无法验证的任务中更可靠地推理。

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

  • 用令牌级反思奖励优化推理质量,聚焦高差异关键步骤。
  • 在四个领域任务中超越基线,提升学习效率与可行性约束满足率。
  • 适合需要可信推理的科研、医疗、法律等高风险场景使用。

在缺乏可验证答案的情况下,对大语言模型进行强化学习训练极具挑战性,即使存在高质量参考答案亦如此。本文提出一种约束型强化学习框架,包含两点:(i) 优化基于链式思维前缀下对参考答案的令牌级密集推理反思奖励(R3),衡量模型对答案的置信度;(ii) 在回溯组层面引入基于评分标准的可行性约束(rubric-gating)。R3 通过识别跨回溯样本间高方差的“推理反射性令牌”,突出那些体现思考差异的关键步骤,避免被大量低方差令牌稀释。该方差信号还用于过滤信号不足的查询,以支持对比学习。Rubric-gating 则将任务的合理准则转化为最终答案的硬性接受/拒绝检查。在涵盖科学写作、医学、法律合同与金融的四个数据集上,该框架显著优于强基线,实现更快、更高效的训练,并有效遵守可行性约束。

原文摘要 · Abstract (English)

Reinforcement learning (RL) training of large language models (LLMs) on unverifiable tasks is challenging even when a reasonable-quality reference answer is available. We propose a constrained RL training framework that (i) optimizes a token-level dense Reasoning Reflection Reward (R3) aligned with reasoning quality, and (ii) enforces rubric-gating as feasibility constraints at the rollout group level. R3 measures the model's token-level certainty of a reference answer under its chain-of-thought (CoT) prefix, and selectively emphasizes tokens with high cross-rollout variance, which we call reasoning-reflective tokens, that would otherwise be diluted by the bulk of low-variance tokens. The same variance signal also drives a filter that discards queries with insufficient signal for comparative learning. Rubric-gating complements R3 by operationalizing principled task criteria as hard accept/reject checks on final answers. Empirically, across four datasets spanning scientific writing, medicine, legal contracts, and finance, our framework outperforms strong baselines, achieves faster, more sample-efficient learning, and respects feasibility constraints.

强化学习推理优化可信生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。