arXiv:2411.04032cs.CL2024-11NAACL被引 29

构建首个专家编辑的机器生成文本基准,评测真实场景下的文本检测效果。

Beemo: Benchmark of Expert-edited Machine-generated Outputs

  • 收集6.5k人类撰写、模型生成及专家编辑的多类文本,覆盖创意写作到摘要等场景。
  • 专家编辑文本能有效规避检测,而模型编辑文本则难被识别为人工撰写。
  • 适合研究文本检测、人机混合内容治理及评估工具的学者与工程师使用。

大语言模型的快速普及增加了机器生成文本(MGT)的数量,并模糊了各类领域中的文本作者身份。然而,现有大多数MGT基准仅包含单作者文本(人类撰写和机器生成),这种设计无法反映更实际的多作者场景——用户对LLM输出进行润色以提升流畅性、连贯性和事实正确性。本文提出基准测试平台Beemo(专家编辑的机器生成文本基准),包含6.5千条由人类撰写、十种指令微调后的LLM生成,以及专家针对不同用途(如创意写作、摘要)编辑的文本;另含13.1千条机器生成并经由模型编辑的文本,支持多种编辑类型下的MGT检测评估。我们详细记录了数据构建流程,并在不同实验设置下对33种MGT检测配置进行了基准测试。结果表明,专家编辑可有效规避检测,而模型编辑文本则难以被识别为人工撰写。Beemo及其全部材料均已公开。

原文摘要 · Abstract (English)

The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most existing MGT benchmarks include single-author texts (human-written and machine-generated). This conventional design fails to capture more practical multi-author scenarios, where the user refines the LLM response for natural flow, coherence, and factual correctness. Our paper introduces the Benchmark of Expert-edited Machine-generated Outputs (Beemo), which includes 6.5k texts written by humans, generated by ten instruction-finetuned LLMs, and edited by experts for various use cases, ranging from creative writing to summarization. Beemo additionally comprises 13.1k machine-generated and LLM-edited texts, allowing for diverse MGT detection evaluation across various edit types. We document Beemo's creation protocol and present the results of benchmarking 33 configurations of MGT detectors in different experimental setups. We find that expert-based editing evades MGT detection, while LLM-edited texts are unlikely to be recognized as human-written. Beemo and all materials are publicly available.

文本检测人机协作基准测试LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。