破解跨语言障碍的书写系统转换技术综述
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP

- 梳理了翻译转写在跨语言NLP中的多种应用动机
- 对比分析不同输入策略的效率与效果差异
- 为研究者提供适配任务需求的实用选择建议
跨语言自然语言处理常受'书写系统障碍'制约,不同文字体系阻碍了语言间知识迁移。转写(transliteration)作为将一种文字系统转换为另一种的技术,能有效提升词汇重叠度,缓解该问题。本文全面综述了转写在跨语言NLP中的应用,提出一个关键动因分类体系,总结了多种将转写作为输入的实现方法。分析了这些方法的演进路径与实际效果,讨论其在现代大模型中的权衡取舍。研究涵盖多种适用场景,包括混合语种文本处理、语言亲缘关系利用以及推理效率优化。基于此,本文为研究者在特定语言、任务和资源条件下选择与实施最合适的转写策略提供了具体建议。
原文摘要 · Abstract (English)
Cross-lingual transfer in NLP is often hindered by the ``script barrier'' where differences in writing systems inhibit transfer learning between languages. Transliteration, the process of converting the script, has emerged as a powerful technique to bridge this gap by increasing lexical overlap. This paper provides a comprehensive survey of the application of transliteration in cross-lingual NLP. We present a taxonomy of key motivations to utilize transliterations in language models, and provide an overview of different approaches of incorporating transliterations as input. We analyze the evolution and effectiveness of these methods, discussing the critical trade-offs involved, and contextualize their need in modern LLMs. The review explores various settings that show how transliteration is beneficial, including handling code-mixed text, leveraging language family relatedness, and pragmatic gains in inference efficiency. Based on this analysis, we provide concrete recommendations for researchers on selecting and implementing the most appropriate transliteration strategy based on their specific language, task, and resource constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。