arXiv:2503.05244cs.AIcs.CL2025-03NeurIPS被引 84

构建首个覆盖6大领域100子领域的写作评估基准,支持动态评分。

WritingBench: A Comprehensive Benchmark for Generative Writing

  • 设计可动态生成评分标准的评估框架,适配不同写作风格。
  • 7B模型经训练后超越GPT-4o在写作任务上的表现。
  • 适合研究文本生成质量、模型评估与写作应用的开发者。

大型语言模型(LLMs)在文本生成方面取得显著进展,但对其生成写作能力的评估仍面临挑战。现有基准多聚焦通用文本生成或有限写作任务,难以全面反映跨领域高质量内容的需求。为此,我们提出WritingBench,一个涵盖6个核心写作领域和100个子领域的综合性评估基准。我们进一步设计了一种查询依赖的评估框架,使LLM能够动态生成特定实例的评估标准。该框架结合微调后的批判模型实现基于风格、格式和长度的感知评分。其有效性通过数据构建能力得到验证:一个7B参数模型在写作任务中表现优于GPT-4o。我们已开源该基准及评估工具与模块化组件,以推动写作类LLM的发展。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have significantly enhanced text generation capabilities, yet evaluating their performance in generative writing remains a challenge. Existing benchmarks primarily focus on generic text generation or limited in writing tasks, failing to capture the diverse requirements of high-quality written contents across various domains. To bridge this gap, we present WritingBench, a comprehensive benchmark designed to evaluate LLMs across 6 core writing domains and 100 subdomains. We further propose a query-dependent evaluation framework that empowers LLMs to dynamically generate instance-specific assessment criteria. This framework is complemented by a fine-tuned critic model for criteria-aware scoring, enabling evaluations in style, format and length. The framework's validity is further demonstrated by its data curation capability, which enables a 7B-parameter model to outperform the performance of GPT-4o in writing. We open-source the benchmark, along with evaluation tools and modular framework components, to advance the development of LLMs in writing.

文本生成评估基准LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。