首份系统综述大模型时代多语言混用技术进展与挑战
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
- 梳理5大领域327篇论文,覆盖80+语言和30+数据集
- 揭示当前大模型在混用语言输入上仍存在性能短板
- 适合关注多语言应用、公平评估的研究者参考
随着大语言模型(LLMs)的快速发展,大多数模型在处理多语言混合输入方面仍存在困难,受限于代码切换(CSW)数据集稀缺及评估偏差,难以在多语言社会中部署。本文首次全面分析了面向代码切换的LLM研究,回顾了涵盖五个研究领域、15个以上NLP任务、30多个数据集和80多个语言的327项研究。按架构、训练策略和评估方法对最新进展进行分类,阐述了大模型如何重塑代码切换建模,并指出现存挑战。文章最后提出路线图,强调需构建包容性数据集、公平评估体系及语言学基础模型,以实现真正多语言能力。
原文摘要 · Abstract (English)
Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual societies. This survey provides the first comprehensive analysis of CSW-aware LLM research, reviewing 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages. We categorize recent advances by architecture, training strategy, and evaluation methodology, outlining how LLMs have reshaped CSW modeling and identifying the challenges that persist. The paper concludes with a roadmap that emphasizes the need for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual capabilities https://github.com/lingo-iitgn/awesome-code-mixing/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。