大模型生成假文本能力提升,但检测系统仍能有效识别。
Cat and Mouse -- Can Fake Text Generation Outpace Detector Systems?
- 用统计分类器检测类侦探小说风格的伪造文本
- Gemini在版本升级后生成更难检测的文本,GPT无明显变化
- 适合关注生成内容安全与检测技术的读者
大型语言模型可在学术写作、产品评论和政治新闻等领域生成逼真的'假文本'。针对此类文本的检测方法已被广泛研究。尽管这看似预示着一场永无止境的'军备竞赛',但我们注意到,新式LLM正使用越来越多的参数、训练数据和能源,而相对简单的分类器却能以少量资源实现良好检测准确率。为探讨模型是否可能最终超越检测系统,我们研究了统计分类器在识别经典侦探小说风格伪造文本方面的能力。在0.5版本迭代中,Gemini展现出更强的欺骗性生成能力,而GPT未见显著提升。这表明即使模型持续增大,可靠检测仍可能保持可行,尽管新型模型架构或可进一步增强其欺骗性。
原文摘要 · Abstract (English)
Large language models can produce convincing "fake text" in domains such as academic writing, product reviews, and political news. Many approaches have been investigated for the detection of artificially generated text. While this may seem to presage an endless "arms race", we note that newer LLMs use ever more parameters, training data, and energy, while relatively simple classifiers demonstrate a good level of detection accuracy with modest resources. To approach the question of whether the models' ability to beat the detectors may therefore reach a plateau, we examine the ability of statistical classifiers to identify "fake text" in the style of classical detective fiction. Over a 0.5 version increase, we found that Gemini showed an increased ability to generate deceptive text, while GPT did not. This suggests that reliable detection of fake text may remain feasible even for ever-larger models, though new model architectures may improve their deceptiveness
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。