arXiv:2506.15683cs.CLcs.CY2025-06被引 1

新方法可检测未见过的私有微调大模型生成文本,准确率超96%

PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning

  • 通过捕捉模型家族共性特征,而非记忆个体差异
  • 在3个模型家族上检测F1超96%,优于7种基线和3个工业服务
  • 适合需要对抗隐蔽私有模型生成内容的场景

随着大语言模型的普及,虚假信息传播和学术不端等问题日益严重,使得检测其生成文本变得前所未有的重要。尽管现有方法已取得显著进展,但由私有微调产生的文本带来的新挑战仍被忽视。用户可通过在开源模型上用私有语料微调,获得私有模型,导致现有检测器性能大幅下降。为此,我们提出PhantomHunter,一种专门用于检测未见过的私有微调大模型生成文本的检测器。其家族感知学习框架聚焦于基础模型及其衍生版本共享的家族级特征,而非记忆个体特征。在LLaMA、Gemma和Mistral三个模型家族的数据上,其表现优于7种基线和3个工业服务,F1分数超过96%。

原文摘要 · Abstract (English)

With the popularity of large language models (LLMs), undesirable societal problems like misinformation production and academic misconduct have been more severe, making LLM-generated text detection now of unprecedented importance. Although existing methods have made remarkable progress, a new challenge posed by text from privately tuned LLMs remains underexplored. Users could easily possess private LLMs by fine-tuning an open-source one with private corpora, resulting in a significant performance drop of existing detectors in practice. To address this issue, we propose PhantomHunter, an LLM-generated text detector specialized for detecting text from unseen, privately-tuned LLMs. Its family-aware learning framework captures family-level traits shared across the base models and their derivatives, instead of memorizing individual characteristics. Experiments on data from LLaMA, Gemma, and Mistral families show its superiority over 7 baselines and 3 industrial services, with F1 scores of over 96%.

文本检测私有微调家族感知大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。