Pangram 4精准识别AI生成文本,误报率仅0.0041%
Pangram 4 Technical Report

- 基于深度学习的文本分类模型,提升细粒度编辑与人机混合文本区分能力
- 在标准基准上实现0.9916的AUROC,误报率0.0041%,漏报率0.3396%
- 对分布外数据和对抗攻击更鲁棒,适合跨领域高精度检测场景
我们介绍Pangram 4,Pangram Labs最新推出的基于深度学习的AI文本分类模型。相比Pangram 3,Pangram 4在整体准确率上有所提升,并展现出更强的分布外泛化能力及对抗攻击鲁棒性。其另一项新贡献是能更好区分细粒度编辑与人机混合撰写内容,显著改善边界检测任务与交错式AI辅助文本的识别效果。在标准AI检测基准上的测试表明,Pangram 4在多种设置和领域下均达到当前最优性能。
原文摘要 · Abstract (English)
We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, Pangram 4 exhibits superior out-of-distribution generalization and robustness to adversarial attacks. Another novel contribution of Pangram 4 is its improved ability to distinguish fine-grained edits and mixed AI-human co-authored text. We demonstrate improvements to both boundary detection tasks and the detection of interleaved AI assistance. Finally, we report metrics on standard AI detection benchmarks showing that Pangram 4 achieves state-of-the-art performance on the AI text detection task across a wide variety of settings and domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。