arXiv:2503.00258cs.CLcs.AI2025-03被引 8

将文本拆解为内容与表达,提升对AI生成文本的检测精度

Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text

  • 把文本分为内容和语言表达两维度,分别分析其异常特征
  • 在两级风险检测中,AUROC最高提升至0.886,显著优于现有方法
  • 适合需要高精度识别AI生成内容的研究者与内容审核团队

大模型的广泛应用对识别文本中AI参与度提出了迫切需求。现有研究多在零散场景下进行,缺乏系统性统一方法。本文提出HART层级风险框架,对应不同检测任务,并设计一种新型二维检测方法,将文本解耦为内容与语言表达。研究发现内容对表面变化具有强鲁棒性,可作为关键检测特征。实验表明,该方法显著优于现有检测器:在二级检测中AUROC从0.705提升至0.849,在RAID数据集上从0.807提升至0.886。代码与数据已开源于https://github.com/baoguangsheng/truth-mirror。

原文摘要 · Abstract (English)

The wide usage of LLMs raises critical requirements on detecting AI participation in texts. Existing studies investigate these detections in scattered contexts, leaving a systematic and unified approach unexplored. In this paper, we present HART, a hierarchical framework of AI risk levels, each corresponding to a detection task. To address these tasks, we propose a novel 2D Detection Method, decoupling a text into content and language expression. Our findings show that content is resistant to surface-level changes, which can serve as a key feature for detection. Experiments demonstrate that 2D method significantly outperforms existing detectors, achieving an AUROC improvement from 0.705 to 0.849 for level-2 detection and from 0.807 to 0.886 for RAID. We release our data and code at https://github.com/baoguangsheng/truth-mirror.

文本检测AI生成深度学习风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。