用音素信息提升大模型对非拉丁语系的语言理解能力
Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
- 用音素转写作为补充信号,让模型捕捉跨文字系统的语音共性
- 在非拉丁文字上性能提升最高达15.1%,拉丁文字也提升12.6%
- 适合需要多语言支持的场景,尤其是低资源非拉丁语系应用
尽管多语言大模型在各类基准上表现优异,但在非拉丁文字语言上仍存在明显短板。原因在于预训练使用以拉丁字符为主的文字系统,掩盖了不同文字间的共同发音特征。本文提出利用音素转写作为补充信号,引导模型学习跨文字的音系表征。实验表明,融合音素信号可显著提升模型在非拉丁及拉丁文字语言上的表现,尤其有效缩小两者间性能差距。详细分析显示,音素与文字分别检索出不同的上下文示例用于提示学习(ICL)。为此我们提出混合式ICL检索策略,结合两者信息,相较随机检索,拉丁文字语言最高提升12.6%,非拉丁文字语言最高提升15.1%。
原文摘要 · Abstract (English)
Although multilingual LLMs have achieved remarkable performance across benchmarks, we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are pretrained with orthographic scripts, which are dominated by Latin characters that obscure their shared phonology with non-Latin scripts. We propose leveraging phonemic transcriptions as complementary signals to induce script-invariant representations. Our study demonstrates that integrating phonemic signals improves performance across both non-Latin and Latin script languages, with a particularly significant impact on closing the performance gap between the two. Through detailed experiments, we show that phonemic and orthographic scripts retrieve distinct examples for in-context learning (ICL). This motivates our proposed Mixed-ICL retrieval strategy, where further aggregation from both leads to our significant performance improvements for both Latin script languages (up to 12.6%) and non-Latin script languages (up to 15.1%) compared to randomized ICL retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。