arXiv:2503.07595cs.CL2025-03被引 2

研究如何绕过大模型生成内容的检测,揭示现有检测工具的脆弱性。

Detection Avoidance Techniques for Large Language Models

  • 通过调节温度、强化学习微调和改写文本实现规避检测。
  • 改写文本可使零样本检测器避过率超90%,原文相似度仍很高。
  • 揭示当前检测系统缺陷,适合安全与内容审核研究者参考。

大型语言模型日益流行,虽推动广泛应用,但也带来传播虚假信息等风险,促使如DetectGPT等分类系统的发展。然而实验表明,这些检测器易受规避技术影响:调节生成模型温度使浅层学习检测器可靠性最低;通过强化学习微调可绕过基于BERT的检测器;改写文本使零样本检测器(如DetectGPT)的避过率超过90%,而文本与原稿高度相似。与现有方法对比显示本文方案性能更优,讨论了其社会影响及未来研究方向。

原文摘要 · Abstract (English)

The increasing popularity of large language models has not only led to widespread use but has also brought various risks, including the potential for systematically spreading fake news. Consequently, the development of classification systems such as DetectGPT has become vital. These detectors are vulnerable to evasion techniques, as demonstrated in an experimental series: Systematic changes of the generative models' temperature proofed shallow learning-detectors to be the least reliable. Fine-tuning the generative model via reinforcement learning circumvented BERT-based-detectors. Finally, rephrasing led to a >90\% evasion of zero-shot-detectors like DetectGPT, although texts stayed highly similar to the original. A comparison with existing work highlights the better performance of the presented methods. Possible implications for society and further research are discussed.

大模型安全检测规避内容伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。