arXiv:2501.13264cs.CL2025-01ACL被引 7

构建开放生成的偏好数据集,提升大模型长文本生成的可靠性。

OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation

  • 设计自动化评估流程,生成33K高质量偏好数据
  • 现有奖励模型在长文本任务表现不佳,新模型显著提升生成质量
  • 适用于长文本生成优化,兼容跨领域数据提升性能

奖励建模在评估和改进大语言模型生成能力中至关重要。尽管已有研究证明其在安全性、帮助性、推理和指令遵循方面的有效性,但在开放式长上下文生成中的能力与泛化性仍鲜有探索。本文提出OpenGenAlign框架及高质量数据集,旨在开发用于评估和改进无幻觉、全面、可靠且高效的开放式长上下文生成的奖励模型。定义四项关键指标,通过自动化流水线在长文本问答、数据转文本和摘要场景中评估多个LLM输出,使用o3完成评估,获得33,000条高质量偏好数据,人类一致率达81%。实验表明,现有奖励模型在保留基准上表现欠佳;而训练的新奖励模型在基准测试中表现更优,并通过强化学习有效提升策略模型的生成质量。此外,OpenGenAlign可用于现有数据集的有效引导生成,且可融合其他领域奖励数据以进一步提升性能。

原文摘要 · Abstract (English)

Reward Modeling is critical in evaluating and improving the generation of Large Language Models (LLMs). While numerous recent works have shown its feasibility in improving safety, helpfulness, reasoning, and instruction-following ability, its capability and generalization to open-ended long-context generation is still rarely explored. In this paper, we introduce OpenGenAlign, a framework and a high-quality dataset designed to develop reward models to evaluate and improve hallucination-free, comprehensive, reliable, and efficient open-ended long-context generation. We define four key metrics to assess generation quality and develop an automated pipeline to evaluate the outputs of multiple LLMs across long-context QA, Data-to-Text, and Summarization scenarios using o3, ending up with 33K high-quality preference data with a human agreement rate of 81\%. Experimental results first demonstrate that existing reward models perform suboptimally on the held-out benchmark. And Our trained reward model achieves superior performance in the benchmark and effectively improves the generation quality of the policy models using Reinforcement Learning (RL). Additionally, OpenGenAlign could be used for effective guided generation in existing datasets. Furthermore, we demonstrate that the OpenGenAlign could be integrated with reward data from other domains to achieve better performance.

奖励建模长文本生成偏好数据集RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。