arXiv:2605.10216cs.CL2026-05

研究发现,语言模型改写会抹去母语痕迹,影响作者母语识别。

The Impact of Editorial Intervention on Detecting Native Language Traces

论文配图:The Impact of Editorial Intervention on Detecting Native Language Traces
图 1 · 摘自论文原文
  • 通过不同级别编辑测试母语识别鲁棒性
  • 轻微修改保留母语特征,流畅改写导致识别率骤降
  • 适合关注AI写作对语言痕迹影响的研究者

母语识别(NLI)旨在从非母语写作中判断作者的母语。随着人机合著兴起,学习者文本常被大模型修正重写,从根本上改变了NLI依赖的语言特征。本文通过在Write & Improve 2024(W&I)语料库的450篇作文上施加不同程度的语法纠错与改写,发现母语归因不只依赖表面错误。检测模型实际上利用更深层的母语相关特征,如非惯用的词汇语义选择和语用迁移。轻微编辑可保留这些结构痕迹并维持高识别准确率;而流畅性优化和改写则使这些母语特征趋于标准化,导致性能显著下降。

原文摘要 · Abstract (English)

Native Language Identification (NLI) is the task of determining an author's native language (L1) from their non-native writing. With the advent of human-AI co-authorship, learner texts are routinely corrected and rewritten by large language models, fundamentally altering the linguistic features NLI approaches depend on. In this paper, we investigate the robustness of L1 traces across increasing degrees of editorial intervention. By processing 450 essays from the Write & Improve 2024 (W&I) corpus through varying levels of grammatical error correction and paraphrasing, we demonstrate that L1 attribution does not depend solely on surface-level errors. Instead, the detection models appear to leverage deeper L1-related features, including unidiomatic lexico-semantic choices and pragmatic transfer. We find that minimal edits preserve these structural traces and maintain high L1 attribution accuracy. In contrast, fluency edits and paraphrasing normalize these L1 features, leading to a sharp decline in performance.

母语识别语言模型文本编辑NLI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。