arXiv:2507.16007cs.CL2025-07ACL被引 14

测试大模型写作文本反馈能力,发现其能给出具体建议但常忽略核心问题。

Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback

  • 构建1300篇带故意写作缺陷的故事数据集,用于评估模型反馈能力。
  • 模型能提供具体且多数准确的反馈,但常错过最严重的问题。
  • 适合对AI辅助创作感兴趣的研究者和内容创作者参考。

大语言模型能否为创意写作者提供有意义的写作反馈?本文通过定义新任务、构建新数据集与评估框架,系统研究模型生成反馈的挑战与局限。为实现可控评估,我们创建了一个包含1,300篇故事的新测试集,其中故意引入各类写作问题。采用自动评估与人工评估相结合的方式,考察主流LLMs在此任务上的表现。分析显示,当前模型在多数情况下展现出较强的开箱即用能力——能提供具体且相对准确的反馈。然而,模型往往无法识别故事中最严重的写作问题,也难以正确判断何时应给出批判性反馈或正面鼓励。

原文摘要 · Abstract (English)

Can LLMs provide support to creative writers by giving meaningful writing feedback? In this paper, we explore the challenges and limitations of model-generated writing feedback by defining a new task, dataset, and evaluation frameworks. To study model performance in a controlled manner, we present a novel test set of 1,300 stories that we corrupted to intentionally introduce writing issues. We study the performance of commonly used LLMs in this task with both automatic and human evaluation metrics. Our analysis shows that current models have strong out-of-the-box behavior in many respects -- providing specific and mostly accurate writing feedback. However, models often fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback.

写作反馈大模型评估提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。