arXiv:2510.12476cs.CLcs.AI2025-10ACL被引 1

发现文本生成检测器在模仿个人风格时会失效,提出预测其失效方向的新方法。

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

  • 识别出个性化文本中特征反转现象导致检测器失准
  • 构建首个个性化生成文本检测基准数据集
  • 新方法可准确预测检测器性能变化,相关性达85%

大型语言模型生成文本愈发流畅,甚至能模仿个人写作风格,但也增加了身份冒用风险。目前尚无研究关注个性化机器生成文本(MGT)的检测。本文提出 extit{dataset},首个用于评估检测器在个性化场景下鲁棒性的基准数据集,基于文学作品与博客原文及其对应的LLM仿写文本构建。实验表明,不同检测器在个性化设置下性能差异显著,部分顶尖模型表现大幅下降。我们归因于特征反转陷阱:通用领域中具有判别力的特征,在个性化文本中出现反向误导。据此提出 extit{method},通过识别潜在的反向特征方向,构建仅沿该方向差异的探测数据集,评估检测器依赖性。实验显示, extit{method}可精准预测性能变化的方向与幅度,与实际性能差距的相关性达85%。本工作旨在推动个性化文本检测研究。

原文摘要 · Abstract (English)

Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this ability also heightens the risk of identity impersonation. To the best of our knowledge, no prior work has examined personalized machine-generated text (MGT) detection. In this paper, we introduce \dataset, the first benchmark for evaluating detector robustness in personalized settings, built from literary and blog texts paired with their LLM-generated imitations. Our experimental results demonstrate large performance gaps across detectors in personalized settings: some state-of-the-art models suffer significant drops. We attribute this limitation to the \textit{feature-inversion trap}, where features that are discriminative in general domains become inverted and misleading when applied to personalized text. Based on this finding, we propose \method, a simple and reliable way to predict detector performance changes in personalized settings. \method identifies latent directions corresponding to inverted features and constructs probe datasets that differ primarily along these features to evaluate detector dependence. Our experiments show that \method can accurately predict both the direction and the magnitude of post-transfer changes, showing 85\% correlation with the actual performance gaps. We hope that this work will encourage further research on personalized text detection.

文本检测特征反转个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。