找出跨模型跨领域最可靠的文本生成特征
A Systematic Analysis of Linguistic Features in AI-Generated Text Detection Across Domains and Models

- 分析284个语言特征在27个模型、10个领域的表现
- 词汇丰富度是唯一跨场景稳定的生成文本信号
- 为非专家提供可解释的生成文本检测方法
可解释的语言特征为说明某段文本为何显得机器生成提供了有前景的途径,尤其对非专业用户而言。然而,现有研究关于哪些特征能可靠指示大语言模型生成文本的结果分散在不同特征集、模型和文本领域之间。为填补这一空白,我们开展了一项大规模实证研究,评估语言信号在表征人工智能生成文本方面的鲁棒性。分析覆盖27个大语言模型和10个文本领域在跨模型与跨领域泛化设置下的284个可解释语言特征。结果表明,仅基于语言特征的分类器能可靠区分人工智能生成与人类写作文本。但许多先前提出的指标表现出强上下文依赖性,唯有词汇丰富度在模型族和文本领域间保持稳健。这些发现揭示了可在不同情境下通用的语言信号,并为更可靠、可解释的人工智能语言分析奠定了基础。
原文摘要 · Abstract (English)
Interpretable linguistic features offer a promising approach for explaining why a given text appears machine-generated, particularly for non-expert users. However, existing findings on which features reliably indicate LLM-generated text remain fragmented across feature sets, models, and text domains. To address this gap, we conduct a large-scale empirical study assessing the robustness of linguistic signals for characterizing AI-generated text. Our analysis covers 284 interpretable linguistic features across outputs from 27 LLMs and ten text domains under cross-model and cross-domain generalization settings. We show that classifiers based solely on linguistic features can reliably distinguish AI-generated from human-written text. However, many previously proposed indicators prove strongly context-dependent, with the exception of measures of lexical richness, which remain robust signals across model families and text domains. These results demonstrate which linguistic signals generalize across contexts and provide a foundation for more reliable, interpretable analyses of AI-generated language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。