arXiv:2603.24981cs.CL2026-03ACL

通过检测隐状态差异,精准放大关键词元,提升文本生成检测鲁棒性。

Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text Detection

  • 基于双模型隐状态差异识别重要词元,实现无训练检测
  • 在DetectRL上平均AUROC提升2.2%,优于现有最优方法
  • 对短文本和局部修改具有强鲁棒性,适合真实场景应用

大语言模型的快速发展正不断模糊人类写作与人工智能生成文本之间的界限,引发虚假信息传播、作者身份模糊及知识产权威胁等社会风险。这凸显了高效可靠检测方法的迫切需求。现有无需训练的方法通常通过聚合词元级信号获得全局得分,但常假设词元贡献均一,在短序列或局部词元修改下表现不佳。为此,我们提出Exons-Detect,一种基于外显子感知词元重加权视角的无训练检测方法。该方法通过双模型设置测量隐状态差异,识别并放大有信息量的外显子词元,再基于重要性加权词元序列计算可解释的翻译得分。实验表明,Exons-Detect达到当前最优检测性能,并对对抗攻击和输入长度变化表现出强鲁棒性,在DetectRL上相较最强基线平均AUROC提升2.2%。

原文摘要 · Abstract (English)

The rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights. These concerns highlight the urgent need for effective and reliable detection methods. While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications. To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective. Exons-Detect identifies and amplifies informative exonic tokens by measuring hidden-state discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence. Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths. In particular, it attains a 2.2\% relative improvement in average AUROC over the strongest prior baseline on DetectRL.

文本检测无训练鲁棒性AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。