提出多检测器融合框架,提升机器生成文本识别能力
Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- 采用多检测器集成方法,统一处理多种文本生成检测任务
- 构建BMAS英文数据集,支持二分类、多分类与生成器溯源
- 覆盖文档级、句子级检测及对抗攻击场景,适用性强
大型语言模型正逼近甚至超越人类创造力,引发对人类工作真实性与创新力保护的关切。本文系统研究机器生成文本检测问题,涵盖文档级二分类与多分类、生成器溯源、句子级边界分割,以及旨在降低可检测性的对抗攻击。为此,提出BMAS英文数据集,支持人类与机器文本的二分类、多类别识别及生成源判定,并能定位人机协作文本的分界点。该工作为机器生成文本检测提供了更全面、实用的解决方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are gearing up to surpass human creativity. The veracity of the statement needs careful consideration. In recent developments, critical questions arise regarding the authenticity of human work and the preservation of their creativity and innovative abilities. This paper investigates such issues. This paper addresses machine-generated text detection across several scenarios, including document-level binary and multiclass classification or generator attribution, sentence-level segmentation to differentiate between human-AI collaborative text, and adversarial attacks aimed at reducing the detectability of machine-generated text. We introduce a new work called BMAS English: an English language dataset for binary classification of human and machine text, for multiclass classification, which not only identifies machine-generated text but can also try to determine its generator, and Adversarial attack addressing where it is a common act for the mitigation of detection, and Sentence-level segmentation, for predicting the boundaries between human and machine-generated text. We believe that this paper will address previous work in Machine-Generated Text Detection (MGTD) in a more meaningful way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。