arXiv:2602.11181cs.CL2026-02

如何让大模型更好理解混合语言代码,这篇指南给出了实用方法。

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

  • 构建混合语言模型需从数据、训练到评估全链路优化
  • 现有评测存在不一致和英语中心偏差问题
  • 适合想提升多语言模型能力的研究者与开发者

代码混用与语言切换(CSW)仍是大语言模型面临的挑战。尽管多语言建模取得进展,现有模型在混合语言场景下仍表现出语法、事实性和安全行为的系统性退化。本文系统梳理现代大模型中的CSW研究,提出涵盖数据、建模与评估维度的统一分类体系,并提炼出可操作的实践指南,涵盖专用于CSW的预训练、任务特定微调、提示工程与上下文学习策略。分析当前评估方法的不稳定性与可复现性问题,盘点现有基准并批判性审视其语言覆盖范围与英语中心偏见。最后讨论新兴安全风险,如利用代码混用绕过模型防护机制,并指出开放研究挑战。

原文摘要 · Abstract (English)

Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling, LLMs often struggle in mixed-language settings, exhibiting systematic degradation in grammaticality, factuality, and safety behavior. This work provides a comprehensive overview of CSW research in modern large language model settings. We introduce a unifying taxonomy that organizes prior work along dimensions of data, modeling, and evaluation, and we distill these findings into a practical playbook of actionable recommendations for building, adapting, and evaluating CSW-capable LLMs. We review modeling approaches ranging from CSW-tailored pre-training and task-specific post-training to prompting strategies and in-context learning. We analyze current evaluation practices, highlighting sources of instability and limited reproducibility, and we catalog existing benchmarks while critically examining their linguistic coverage and English-centric biases. Finally, we discuss emerging safety concerns, including use of code-mixing as a mechanism for bypassing model safeguards, and identify open research challenges.

多语言代码混用大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。