arXiv:2501.16836cs.CLcs.AI2025-01综述被引 3

系统梳理了自然语言处理中拼写错误的挑战与解决方案。

Misspellings in Natural Language Processing: A survey

  • 从数据增强到字符顺序无关等方法,应对拼写错误影响。
  • 揭示拼写错误对文本分类、机器翻译等任务的性能下降问题。
  • 适合关注NLP鲁棒性、安全伦理及大模型测试的研究者。

本综述全面探讨了自然语言处理中拼写错误带来的挑战。尽管拼写错误在数字通信中普遍存在,尤其在社交媒体、博客和论坛等用户生成内容场景中,人类通常能理解其含义,但现有NLP模型常因无法有效处理而性能下降。本文回顾了拼写错误作为科学问题的发展历程,并总结了最新应对策略,包括数据增强、双步处理、字符顺序无关、基于元组的方法等。同时,文章分析了相关数据集、竞赛及其对领域推动的作用,讨论了拼写错误被用于传播恶意信息和仇恨言论等伦理安全风险。此外,还从心理语言学角度解析人类如何处理拼写错误,为文本归一化与表征提供新思路。最后,评估了主流大语言模型在拼写错误下的表现与基准测试现状。该综述旨在为研究者提供全面资源,以缓解拼写错误对不断演进的NLP环境的影响。

原文摘要 · Abstract (English)

This survey provides an overview of the challenges of misspellings in natural language processing (NLP). While often unintentional, misspellings have become ubiquitous in digital communication, especially with the proliferation of Web 2.0, user-generated content, and informal text mediums such as social media, blogs, and forums. Even if humans can generally interpret misspelled text, NLP models frequently struggle to handle it: this causes a decline in performance in common tasks like text classification and machine translation. In this paper, we reconstruct a history of misspellings as a scientific problem. We then discuss the latest advancements to address the challenge of misspellings in NLP. Main strategies to mitigate the effect of misspellings include data augmentation, double step, character-order agnostic, and tuple-based methods, among others. This survey also examines dedicated data challenges and competitions to spur progress in the field. Critical safety and ethical concerns are also examined, for example, the voluntary use of misspellings to inject malicious messages and hate speech on social networks. Furthermore, the survey explores psycholinguistic perspectives on how humans process misspellings, potentially informing innovative computational techniques for text normalization and representation. Finally, the misspelling-related challenges and opportunities associated with modern large language models are also analyzed, including benchmarks, datasets, and performances of the most prominent language models against misspellings. This survey aims to be an exhaustive resource for researchers seeking to mitigate the impact of misspellings in the rapidly evolving landscape of NLP.

拼写纠错NLP鲁棒性大模型评估安全伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。