通过分析事件演变模式,区分机器与人类长文本。
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
- 用事件序列建模文本潜在空间,避开词汇层面干扰。
- 在三个领域中检测准确率提升31%,优于DetectGPT等基线。
- 揭示大模型事件触发方式与人类本质差异,适合内容审核场景。
大型语言模型生成流畅长文本的能力日益增强,给区分机器生成内容与人类写作带来挑战,这对保障表达的真实性与可信度至关重要。现有零样本检测方法主要依赖词元级分布,易受真实世界领域偏移影响,包括不同提示和解码策略,以及对抗攻击。本文提出一种更鲁棒的方法:将事件过渡等抽象元素作为关键判别因素,通过在人类写作的事件或主题序列上训练潜在空间模型来检测机器与人类文本。在三个不同领域中,原本在词元层面难以区分的机器生成文本,可通过该潜在空间模型更好识别,相比DetectGPT等强基线实现31%的性能提升。进一步分析表明,现代大模型如GPT-4在事件触发及其过渡方式上与人类存在本质差异,这一内在差别使本方法能稳健检测机器生成文本。
原文摘要 · Abstract (English)
The increasing capability of large language models (LLMs) to generate fluent long-form texts is presenting new challenges in distinguishing machine-generated outputs from human-written ones, which is crucial for ensuring authenticity and trustworthiness of expressions. Existing zero-shot detectors primarily focus on token-level distributions, which are vulnerable to real-world domain shifts, including different prompting and decoding strategies, and adversarial attacks. We propose a more robust method that incorporates abstract elements, such as event transitions, as key deciding factors to detect machine versus human texts by training a latent-space model on sequences of events or topics derived from human-written texts. In three different domains, machine-generated texts, which are originally inseparable from human texts on the token level, can be better distinguished with our latent-space model, leading to a 31% improvement over strong baselines such as DetectGPT. Our analysis further reveals that, unlike humans, modern LLMs like GPT-4 generate event triggers and their transitions differently, an inherent disparity that helps our method to robustly detect machine-generated texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。