arXiv:2502.12611cs.CL2025-02ACL被引 5

研究作者背景如何影响AI文本检测准确率,发现语言水平和环境是关键因素。

Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection

  • 分析性别、语言水平等作者特征对检测器的影响
  • 语言水平和语言环境显著降低检测准确率
  • 为公平可靠的AI检测系统提供实证依据

大语言模型(LLMs)的兴起要求精准的AI生成文本检测。然而,现有方法普遍忽视作者特征的影响。本文研究了社会语言学属性——性别、CEFR语言水平、学术领域和语言环境——对先进AI文本检测器的影响。基于ICNALE人类撰写文本语料库及多种LLM生成的平行文本,采用多因素方差分析和加权最小二乘法(WLS)进行严格评估。结果表明:语言水平和语言环境对检测准确率有显著影响;性别与学术领域的影响则依赖于具体检测器。这些发现凸显了在真实、跨域场景中发展社会敏感型检测系统的必要性,以避免对特定群体产生不公平判罚。本文提供了新的实证证据、稳健的统计框架和可操作洞见,推动未来在偏差缓解、包容性评估基准和社会责任导向的LLM检测器方面的研究。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) necessitates accurate AI-generated text detection. However, current approaches largely overlook the influence of author characteristics. We investigate how sociolinguistic attributes-gender, CEFR proficiency, academic field, and language environment-impact state-of-the-art AI text detectors. Using the ICNALE corpus of human-authored texts and parallel AI-generated texts from diverse LLMs, we conduct a rigorous evaluation employing multi-factor ANOVA and weighted least squares (WLS). Our results reveal significant biases: CEFR proficiency and language environment consistently affected detector accuracy, while gender and academic field showed detector-dependent effects. These findings highlight the crucial need for socially aware AI text detection to avoid unfairly penalizing specific demographic groups. We offer novel empirical evidence, a robust statistical framework, and actionable insights for developing more equitable and reliable detection systems in real-world, out-of-domain contexts. This work paves the way for future research on bias mitigation, inclusive evaluation benchmarks, and socially responsible LLM detectors.

AI检测社会偏见语言水平公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。