arXiv:2409.14509cs.CLcs.CY2024-09被引 68

通过专业编辑发现AI写作缺陷并改进人机写作风格对齐

Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits

  • 构建七类问题分类体系,系统归纳AI文本常见弊端
  • 创建1057段编辑数据集LAMP,验证主流模型无显著优劣
  • 自动编辑方法虽不及人工,但显著提升文本人类感

基于大语言模型的应用正广泛用于写作,生成内容已进入社交媒体、新闻和课堂。然而,其与人类写作的差异仍不清晰。我们邀请专业写作者在多个创意领域对段落进行编辑。首先,发现写作者对LLM生成文本中的不良特征达成共识,并将其归纳为七类(如陈词滥调、冗余说明)。其次,构建了包含1057段由专业写作者根据该分类体系编辑的LLM生成段落的LAMP数据集。分析显示,本研究涉及的各模型(GPT4o、Claude-3.5-Sonnet、Llama-3.1-70b)在写作质量上未分高下,表明跨模型存在共性局限。第三,基于已有自动编辑工作评估改进方法。大规模偏好标注证实,尽管专家仍更倾向人工编辑文本,但自动编辑方法在提升生成文本与人类写作风格对齐方面展现出良好前景。

原文摘要 · Abstract (English)

LLM-based applications are helping people write, and LLM-generated text is making its way into social media, journalism, and our classrooms. However, the differences between LLM-generated and human written text remain unclear. To explore this, we hired professional writers to edit paragraphs in several creative domains. We first found these writers agree on undesirable idiosyncrasies in LLM generated text, formalizing it into a seven-category taxonomy (e.g. clichés, unnecessary exposition). Second, we curated the LAMP corpus: 1,057 LLM-generated paragraphs edited by professional writers according to our taxonomy. Analysis of LAMP reveals that none of the LLMs used in our study (GPT4o, Claude-3.5-Sonnet, Llama-3.1-70b) outperform each other in terms of writing quality, revealing common limitations across model families. Third, building on existing work in automatic editing we evaluated methods to improve LLM-generated text. A large-scale preference annotation confirms that although experts largely prefer text edited by other experts, automatic editing methods show promise in improving alignment between LLM-generated and human-written text.

AI写作风格对齐文本编辑人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。