arXiv:2503.07990cs.CL2025-03被引 4

用混语数据微调mBERT,提升多语言模型在混合语境下的表现

Enhancing Multilingual Language Models for Code-Switched Input Data

  • 在斯帕英推文数据上微调mBERT,增强混语处理能力
  • 语法标注任务性能显著提升,其他任务持平或优于基线
  • 发现英语西班牙语嵌入更一致,适合研究跨语言交互

代码切换(在同一对话中交替使用语言)对自然语言处理任务中的多语言语言模型构成挑战。本研究探讨在混语数据上预训练多语言BERT(mBERT)是否能提升其在词性标注、情感分析、命名实体识别和语言识别等关键任务上的表现。实验采用斯帕英推文数据集进行预训练,并与基线模型对比。结果表明,所提预训练mBERT模型在各项任务中表现优于或匹配基线,尤其在词性标注任务中提升显著。此外,隐层分析显示语言识别任务中英语与西班牙语嵌入更为同质,为未来建模提供启示。研究证明,通过适应混语输入,可增强多语言模型在全球化多语言场景下的实用性。未来工作包括拓展至其他语对、融合多形式数据,以及探索更优的上下文依赖型代码切换理解方法。

原文摘要 · Abstract (English)

Code-switching, or alternating between languages within a single conversation, presents challenges for multilingual language models on NLP tasks. This research investigates if pre-training Multilingual BERT (mBERT) on code-switched datasets improves the model's performance on critical NLP tasks such as part of speech tagging, sentiment analysis, named entity recognition, and language identification. We use a dataset of Spanglish tweets for pre-training and evaluate the pre-trained model against a baseline model. Our findings show that our pre-trained mBERT model outperforms or matches the baseline model in the given tasks, with the most significant improvements seen for parts of speech tagging. Additionally, our latent analysis uncovers more homogenous English and Spanish embeddings for language identification tasks, providing insights for future modeling work. This research highlights potential for adapting multilingual LMs for code-switched input data in order for advanced utility in globalized and multilingual contexts. Future work includes extending experiments to other language pairs, incorporating multiform data, and exploring methods for better understanding context-dependent code-switches.

多语言模型代码切换嵌入分析mBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。