用编辑奖励和测试时计算提升AI写作质量,让机器更懂好文标准。
AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
- 构建写作质量评估基准WQ,整合5个数据集共4729条判断
- 训练的奖励模型在多个测试集上达74%准确率,显著优于基线
- 利用额外算力生成并筛选改写稿,专家偏好率达72.2%
AI生成文本已广泛应用于创意写作、新闻报道、营销内容及科学文章等领域。尽管模型能生成连贯且语法正确的输出,但如何评估和提升其写作质量仍缺乏研究。本文首次整合五个写作偏好数据集,构建包含4729条判断的写作质量基准(WQ)。实验表明,多数先进大模型在该基准上的表现仅略高于随机基线。为此,我们训练了多种规模的写作质量奖励模型(WQRM),在四个分布外测试集上展现出强泛化能力,且在WQ基准上达到74%准确率。进一步地,通过引入测试时计算,生成并排序多个候选改写版本,从而从初稿中选出更高质量内容。九位经验丰富的写作者参与的人类评估显示,基于WQRM的选稿整体获得66%的专家青睐,当奖励差距超过1分时,这一比例升至72.2%。论文公开了数据集与模型,旨在推动社区对写作质量评估的关注与研究。
原文摘要 · Abstract (English)
AI-generated text is proliferating across domains, from creative writing and journalism to marketing content and scientific articles. Models can follow user-provided instructions to generate coherent and grammatically correct outputs but in this work, we study a more fundamental question: how do we evaluate and improve the writing quality of AI-generated text? Writing quality assessment has received less attention from the community, in part because it is fundamentally subjective and requires expertise. We first introduce the Writing Quality Benchmark (WQ) by consolidating five writing-preference datasets into 4,729 writing quality judgments. Our experiments show that most of the competitive baselines, including state-of-the-art LLMs that excel at reasoning tasks, barely outperform random baselines on WQ. We then train specialized Writing Quality Reward Models (WQRM) of various sizes for writing quality assessment that demonstrate strong generalization on four out-of-distribution test sets and 74% accuracy on the WQ benchmark. To further show WQRM's practical benefits during inference, we leverage additional test-time compute to generate and rank multiple candidate revisions, allowing us to select higher-quality outputs from an initial draft. Human evaluation with 9 experienced writers confirm that WQRM-based selection produces writing samples preferred by experts 66% overall, and 72.2% when the reward gap is larger than 1 point. We release our datasets and models to encourage community engagement with writing quality assessment and development of AI writing systems better aligned with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。