用解耦奖励机制让大模型通用改写更精准,兼顾事实、风格和对话场景。
Dr Genre: Reinforcement Learning from Decoupled LLM Feedback for Generic Text Rewriting
- 拆分奖励信号,按任务需求动态调整不同目标权重。
- 在多个数据集上提升指令遵循、连贯性和简洁性表现。
- 适合需要多场景文本改写的开发者与研究者使用。
通用文本改写是大语言模型的常见应用,涵盖风格迁移、事实修正和邮件编辑等多样任务。这些任务目标各异(如事实一致性与语义保留),难以用单一模型全面胜任。现有方法通常仅针对特定任务或目标,泛化能力受限。本文提出一种通用改写模型,结合自建对话式改写数据集ChatRewrite,以及LongFact(事实改写)和RewriteLM(风格改写)等主流数据集,构建全面评估基准。为适配不同任务目标,提出Dr Genre——一种解耦奖励学习框架,采用面向目标的奖励模型并引入任务特定权重。实验表明,该方法在指令遵循(一致性)、内部一致性和简洁性等指标上均优于基线,全面提升各类改写质量。
原文摘要 · Abstract (English)
Generic text rewriting is a prevalent large language model (LLM) application that covers diverse real-world tasks, such as style transfer, fact correction, and email editing. These tasks vary in rewriting objectives (e.g., factual consistency vs. semantic preservation), making it challenging to develop a unified model that excels across all dimensions. Existing methods often specialize in either a single task or a specific objective, limiting their generalizability. In this work, we introduce a generic model proficient in factuality, stylistic, and conversational rewriting tasks. To simulate real-world user rewrite requests, we construct a conversational rewrite dataset, ChatRewrite, that presents ``natural''-sounding instructions, from raw emails using LLMs. Combined with other popular rewrite datasets, including LongFact for the factuality rewrite task and RewriteLM for the stylistic rewrite task, this forms a broad benchmark for training and evaluating generic rewrite models. To align with task-specific objectives, we propose Dr Genre, a Decoupled-reward learning framework for Generic rewriting, that utilizes objective-oriented reward models with a task-specific weighting. Evaluation shows that \approach delivers higher-quality rewrites across all targeted tasks, improving objectives including instruction following (agreement), internal consistency (coherence), and minimal unnecessary edits (conciseness).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。