通过语言特征分析,揭示AI文本检测器泛化失败的深层原因。
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
- 构建涵盖6种提示、7个模型、4个领域的综合数据集
- 发现时态使用和代词频率等特征与泛化性能强相关
- 为检测器改进提供可解释的优化方向,适合检测研究者参考
AI文本检测器在特定领域基准上表现优异,但在跨提示、跨模型或跨领域时泛化能力差。现有研究虽指出这一现象,但缺乏成因分析。本文系统性地通过语言学分析解释泛化行为:构建包含6种提示策略、7个大语言模型(LLMs)和4个领域数据集的综合性基准,生成多样化的真人与AI生成文本;在不同生成条件下微调分类型检测器,并评估其跨提示、跨模型、跨数据集的泛化能力;进一步计算80种语言学特征在训练与测试条件间的特征偏移与泛化准确率的相关性。结果表明,特定检测器在特定评估场景下的泛化性能显著关联于时态使用、代词频率等语言特征。
原文摘要 · Abstract (English)
AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or domains. While prior work has reported these generalization gaps, there are limited insights about the underlying causes. In this work, we present a systematic study aimed at explaining generalization behavior through linguistic analysis. We construct a comprehensive benchmark that spans 6 prompting strategies, 7 large language models (LLMs), and 4 domain datasets, resulting in a diverse set of human- and AI-generated texts. Using this dataset, we fine-tune classification-based detectors on various generation settings and evaluate their cross-prompt, cross-model, and cross-dataset generalization. To explain the performance variance, we compute correlations between generalization accuracies and feature shifts of 80 linguistic features between training and test conditions. Our analysis reveals that generalization performance for specific detectors and evaluation conditions is significantly associated with linguistic features such as tense usage and pronoun frequency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。