arXiv:2602.22145cs.HCcs.AI2026-02

AI正悄悄抹去非英语母语者的语言特色,这篇论文量化了这一现象。

When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models

  • 提出‘文化幽灵化’概念,用新指标衡量AI对非母语英语特征的消除
  • 发现模型平均擦除10.26%文化标记,礼貌用语比词汇更易被删改
  • 明确提示保留文化特征可降30%擦除率,且不影响内容质量

大型语言模型(LLMs)在职场沟通中被广泛用于“专业化”文本,却常以牺牲语言身份为代价。本文提出“文化幽灵化”(Cultural Ghosting),指非母语英语变体中独特语言标记在文本处理过程中的系统性消失。通过分析5个模型在三种提示条件下对1,490篇具有文化标记的文本(印度英语、新加坡英语、尼日利亚英语)生成的22,350条输出,我们引入两个新指标:身份擦除率(IER)与语义保留得分(SPS)。所有提示下总体IER为10.26%,模型间差异达3.5%至20.5%(5.9倍范围)。关键发现:尽管语义相似度高(均值SPS=0.748),但文化标记仍被系统性删除。实用标记(礼貌惯例)擦除率(71.5%)是词汇标记(37.1%)的1.9倍。实验表明,加入明确的文化保留提示可使擦除率降低29%,同时不损害语义质量。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to ``professionalize'' workplace communication, often at the cost of linguistic identity. We introduce "Cultural Ghosting", the systematic erasure of linguistic markers unique to non-native English varieties during text processing. Through analysis of 22,350 LLM outputs generated from 1,490 culturally marked texts (Indian, Singaporean,& Nigerian English) processed by five models under three prompt conditions, we quantify this phenomenon using two novel metrics: Identity Erasure Rate (IER) & Semantic Preservation Score (SPS). Across all prompts, we find an overall IER of 10.26%, with model-level variation from 3.5% to 20.5% (5.9x range). Crucially, we identify a Semantic Preservation Paradox: models maintain high semantic similarity (mean SPS = 0.748) while systematically erasing cultural markers. Pragmatic markers (politeness conventions) are 1.9x more vulnerable than lexical markers (71.5% vs. 37.1% erasure). Our experiments demonstrate that explicit cultural-preservation prompts reduce erasure by 29% without sacrificing semantic quality.

语言模型文化偏见语用标记AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。