arXiv:2502.07424cs.CLcs.AI2025-02ACL被引 11

发现大模型处理非罗马字母语言时会隐式使用罗马化表征

RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs

  • 通过可解释性分析发现中间层先用罗马字母表示非罗马文字
  • 罗马化与原文字在语义表征上高度相似,存在共享空间
  • 翻译目标为罗马化形式时,模型更早激活对应表征

大型语言模型虽主要在英语数据上训练,却表现出强大的多语言能力。针对使用非罗马字母语言的情况,本文研究罗马化(用罗马字母表示非罗马文字)是否在多语言处理中充当桥梁。通过机制可解释性技术分析下一个词生成过程,发现中间层常以罗马化形式表示目标词,随后才转为原生文字,这一现象称为隐式罗马化。进一步的激活修补实验表明,模型对原生文本与罗马化文本的语义概念编码方式相似,说明存在共享的底层表征。此外,在将内容翻译为非罗马字母语言时,若目标语言以罗马化形式呈现,其表征会在模型更早的层中出现。这些发现深化了对大模型多语言表征的理解,揭示了罗马化在语言迁移中的隐性作用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit strong multilingual performance despite being predominantly trained on English-centric corpora. This raises a fundamental question: How do LLMs achieve such multilingual capabilities? Focusing on languages written in non-Roman scripts, we investigate the role of Romanization - the representation of non-Roman scripts using Roman characters - as a potential bridge in multilingual processing. Using mechanistic interpretability techniques, we analyze next-token generation and find that intermediate layers frequently represent target words in Romanized form before transitioning to native script, a phenomenon we term Latent Romanization. Further, through activation patching experiments, we demonstrate that LLMs encode semantic concepts similarly across native and Romanized scripts, suggesting a shared underlying representation. Additionally, for translation into non-Roman script languages, our findings reveal that when the target language is in Romanized form, its representations emerge earlier in the model's layers compared to native script. These insights contribute to a deeper understanding of multilingual representation in LLMs and highlight the implicit role of Romanization in facilitating language transfer.

大模型多语言可解释性罗马化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。