arXiv:2601.16512cs.CL2026-01Conference of the …被引 1

用搜索引擎找原文,识别大模型改写过的文本。

SearchLLM: Detecting LLM Paraphrased Text by Measuring the Similarity with Regeneration of the Candidate Source via Search Engine

  • 通过搜索潜在原文并比对重生成内容,检测改写文本。
  • 在多个大模型上提升检测准确率,有效防伪改。
  • 可嵌入现有检测器,适合反抄袭与内容安全场景。

随着大语言模型(LLMs)的普及,用户常利用其进行文本润色和改写,但可能导致原意丢失或扭曲。由于大模型生成内容接近人类写作,传统检测方法难以识别高度模仿原文的改写文本。为此,我们提出SearchLLM,一种基于搜索引擎定位潜在原始来源的新方法。通过分析输入文本与候选源文本重生成版本之间的相似性,能有效区分大模型改写内容。SearchLLM作为代理层,可无缝集成至现有检测器以提升性能。实验表明,该方法在多种大模型生成的近似原文改写文本检测中均显著提高准确率,并有助于防御改写攻击。

原文摘要 · Abstract (English)

With the advent of large language models (LLMs), it has become common practice for users to draft text and utilize LLMs to enhance its quality through paraphrasing. However, this process can sometimes result in the loss or distortion of the original intended meaning. Due to the human-like quality of LLM-generated text, traditional detection methods often fail, particularly when text is paraphrased to closely mimic original content. In response to these challenges, we propose a novel approach named SearchLLM, designed to identify LLM-paraphrased text by leveraging search engine capabilities to locate potential original text sources. By analyzing similarities between the input and regenerated versions of candidate sources, SearchLLM effectively distinguishes LLM-paraphrased content. SearchLLM is designed as a proxy layer, allowing seamless integration with existing detectors to enhance their performance. Experimental results across various LLMs demonstrate that SearchLLM consistently enhances the accuracy of recent detectors in detecting LLM-paraphrased text that closely mimics original content. Furthermore, SearchLLM also helps the detectors prevent paraphrasing attacks.

文本检测大模型改写识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。