研究英语母语者与打字错误对大模型性能的联合影响。
Individual and Combined Effects of English as a Second Language and Typos on LLM Performance
- 用Trans-EnV生成8种非母语英语变体,用MulTypo注入三档错别字。
- 二者叠加导致性能下降更严重,闭合任务中效果最明显。
- 提醒评估应考虑真实使用场景,避免高估模型实际表现。
大型语言模型(LLMs)在全球广泛应用,其训练数据主要为英语,因此在英文输入下表现最佳。许多非英语母语者以第二语言(ESL)使用这些模型,且输入常含拼写错误。以往研究多分别考察ESL变化与拼写错误,但两者在实际使用中常同时出现。本研究采用Trans-EnV框架将标准英语输入转化为8种ESL变体,并利用MulTypo在三个水平(低、中、高)注入拼写错误。结果表明,ESL与拼写错误的联合效应通常比单一因素导致的性能下降更显著,且非简单叠加。该趋势在闭合式任务中最为明显,性能退化可更一致地刻画;而开放式任务结果则更混杂。总体而言,仅在标准英语上评估可能过高估计模型真实性能,单独评估ESL或拼写错误无法充分反映真实场景下的模型行为。
原文摘要 · Abstract (English)
Large language models (LLMs) are used globally, and because much of their training data is in English, they typically perform best on English inputs. As a result, many non-native English speakers interact with them in English as a second language (ESL), and these inputs often contain typographical errors. Prior work has largely studied the effects of ESL variation and typographical errors separately, even though they often co-occur in real-world use. In this study, we use the Trans-EnV framework to transform standard English inputs into eight ESL variants and apply MulTypo to inject typos at three levels: low, moderate, and severe. We find that combining ESL variation and typos generally leads to larger performance drops than either factor alone, though the combined effect is not simply additive. This pattern is clearest on closed-ended tasks, where performance degradation can be characterized more consistently across ESL variants and typo levels, while results on open-ended tasks are more mixed. Overall, these findings suggest that evaluations on clean standard English may overestimate real-world model performance, and that evaluating ESL variation and typographical errors in isolation does not fully capture model behavior in realistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。