arXiv:2502.11614cs.CLcs.AI2025-02ACL被引 3

人类能识别AI生成文本,且偏好不总是偏向真人写作。

Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI

  • 跨9语言9领域16数据集验证人类检测能力上限
  • 平均识别准确率达87.6%,远超随机水平
  • 提示显式说明差异可提升50%以上检测效果

以往研究认为人类难以区分大语言模型生成文本与真人写作,常接近随机猜测。为验证这一结论在多语言、多领域的普适性,我们开展大规模案例研究,评估人类检测的上限。在覆盖9种语言和9个领域的16个数据集上,19名标注者平均检测准确率达87.6%,挑战了既有结论。研究发现,人机文本的主要差异在于具体性、文化细节和多样性;通过在提示中明确解释这些差异,可在超过50%的情况下部分弥合差距。但同时发现,当人类无法清晰判断来源时,反而不总是偏好真人文本。相关数据集、人工标签及标注者元数据已公开于https://github.com/xnlp-lab/HumanEval-MGT。

原文摘要 · Abstract (English)

Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random guessing. To verify the generalizability of this finding across languages and domains, we perform an extensive case study to identify the upper bound of human detection accuracy. Across 16 datasets covering 9 languages and 9 domains, 19 annotators achieved an average detection accuracy of 87.6%, thus challenging previous conclusions. We find that major gaps between human and machine text lie in concreteness, cultural nuances, and diversity. Prompting by explicitly explaining the distinctions in the prompts can partially bridge the gaps in over 50% of the cases. However, we also find that humans do not always prefer human-written text, particularly when they cannot clearly identify its source. We release our dataset, the human labels, and the annotator metadata at https://github.com/xnlp-lab/HumanEval-MGT.

文本检测多语言人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。