arXiv:2506.16024cs.CLcs.AI2025-06被引 4

用智能奖励机制提升大模型长文本生成能力,超越GPT-4

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation

  • 基于强化学习设计靶向奖励信号,替代通用评估
  • 自动构建数据集,训练后性能提升20%且超GPT-4-Turbo
  • 适合需要精准长文生成的复杂问答场景

当前大语言模型在长上下文理解方面研究较多,但开放式长文本生成(Open-LTG)仍不充分。训练长文本生成模型需高质量参考数据,而此类数据在信息性任务中通常缺失。现有方法仅使用通用评估作为奖励信号,限制了生成精度。为此,我们提出ProxyReward框架,包含数据集构建与奖励信号计算方法:首先通过简单提示自动生成数据集,避免大量人工标注;其次提供针对特定问题的信息完整性和准确性靶向评估。实验表明,ProxyReward在开放长文本生成任务中表现优于GPT-4-Turbo,可使主流开源模型性能提升20%,同时超越以大模型为裁判的方法。该工作有效提升了大模型应对复杂人类开放式问题的能力。

原文摘要 · Abstract (English)

Current research on long-form context in Large Language Models (LLMs) primarily focuses on the understanding of long-contexts, the Open-ended Long Text Generation (Open-LTG) remains insufficiently explored. Training a long-context generation model requires curation of gold standard reference data, which is typically nonexistent for informative Open-LTG tasks. However, previous methods only utilize general assessments as reward signals, which limits accuracy. To bridge this gap, we introduce ProxyReward, an innovative reinforcement learning (RL) based framework, which includes a dataset and a reward signal computation method. Firstly, ProxyReward Dataset generation is accomplished through simple prompts that enables the model to create automatically, obviating extensive labeled data or significant manual effort. Secondly, ProxyReward Signal offers a targeted evaluation of information comprehensiveness and accuracy for specific questions. The experimental results indicate that our method ProxyReward surpasses even GPT-4-Turbo. It can significantly enhance performance by 20% on the Open-LTG task when training widely used open-source models, while also surpassing the LLM-as-a-Judge approach. Our work presents effective methods to enhance the ability of LLMs to address complex open-ended questions posed by human.

长文本生成强化学习大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。