警告可能被检测后,人写AI辅助文章更像真人,但文本特征无差异。
Can Humans Detect AI? Mining Textual Signals of AI-Assisted Writing Under Varying Scrutiny Conditions

- 实验分两组:一组被告知会被检测,另一组不知情,均用同一AI工具写作。
- 54.13%的评委认为被警告组的文章更像人类写的,显著高于随机水平。
- 尽管文本特征(如词汇多样性、句式)完全相同,评委仍能感知差异。
本研究探讨了被检测威胁是否影响人们使用AI写作的方式,以及他人能否识别出差异。在双阶段受控实验中,21名参与者使用AI聊天机器人撰写关于远程工作的观点文章。一半参与者被随机告知其提交内容将被AI检测工具扫描,另一半未被告知。两组均可访问相同的聊天机器人。第二阶段中,251名独立评审对1999组配对文档进行评估,每次判断哪一篇更可能是人类所写。评审者未被告知双方均使用了AI。总体上,评审选择被告知组文档为人类写作的比例为54.13%,未被告知组为45.87%。双尾二项检验显示结果显著偏离随机猜测(p = 0.000243),且该趋势在两种写作立场下均成立。然而,在所有可测量的文本特征(包括AI重叠度、词汇多样性、句式结构和代词使用)上,两组表现无差异。评审的判断依赖于特征分析无法捕捉的潜在信号。
原文摘要 · Abstract (English)
This study asks whether the threat of AI detection changes how people write with AI, and whether other people can tell the difference. In a two-phase controlled experiment, 21 participants wrote opinion pieces on remote work using an AI chatbot. Half were randomly warned that their submission would be scanned by an AI detection tool. The other half received no warning. Both groups had access to the same chatbot. In Phase 2, 251 independent judges evaluated 1,999 paired comparisons, each time choosing which document in the pair was written by a human. Judges were not told that both writers had access to AI. Across all evaluations, judges selected the warned writer's document as human 54.13% of the time versus 45.87% for the unwarned writer. A two-sided binomial test rejects chance guessing at p = 0.000243, and the result holds across both writing stances. Yet on every measurable text feature extracted, including AI overlap scores, lexical diversity, sentence structure, and pronoun usage, the two groups were indistinguishable. The judges are picking up on something that feature-based methods do not capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。