针对混合生成文本的精准检测,支持多语言、跨模型、抗干扰。
Robust and Fine-Grained Detection of AI Generated Texts
- 基于令牌分类构建模型,专用于识别人类与大模型协作生成的内容。
- 在23种语言、未见模型、非母语者文本上均保持高准确率。
- 适用于真实场景中复杂混合文本,尤其适合内容审核与学术诚信检测。
理想的机器生成内容检测系统应能应对不断涌现的各类先进大模型。现有系统在短文本检测上表现不佳,且多数忽略人类与大模型共同创作的混合文本。本文提出一套用于令牌分类的检测模型,训练于超过240万条人类与机器协作生成的文本,覆盖23种语言,涵盖多个主流商用大模型。模型在未见过的领域、未见过的生成器、非母语文本及对抗性输入下表现稳健。我们还构建了新的大规模数据集,并分析了模型在不同领域、生成器、输入长度和对抗方法下的性能差异,揭示了生成文本与原始人工文本在特征上的关键区别。
原文摘要 · Abstract (English)
An ideal detection system for machine generated content is supposed to work well on any generator as many more advanced LLMs come into existence day by day. Existing systems often struggle with accurately identifying AI-generated content over shorter texts. Further, not all texts might be entirely authored by a human or LLM, hence we focused more over partial cases i.e human-LLM co-authored texts. Our paper introduces a set of models built for the task of token classification which are trained on an extensive collection of human-machine co-authored texts, which performed well over texts of unseen domains, unseen generators, texts by non-native speakers and those with adversarial inputs. We also introduce a new dataset of over 2.4M such texts mostly co-authored by several popular proprietary LLMs over 23 languages. We also present findings of our models' performance over each texts of each domain and generator. Additional findings include comparison of performance against each adversarial method, length of input texts and characteristics of generated texts compared to the original human authored texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。