多语言大模型难区分跨语言同形词的语义,依赖拼写而非意义理解。
Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing
- 测试模型在孤立词与句子中处理同形词、类比词和跨语言同形词的能力。
- 跨语言同形词歧义消解准确率低于随机水平,拼写相似性主导判断。
- 模型对语义理解能力差,且对英语与非英语同形词策略不一致。
双语词汇处理受音位、拼写和语义特征的复杂交互影响。人类能轻松处理同形词(如英语和德语的'blind'均意为'失明'),而跨语言同形词(如英语'gift'为'礼物',德语'gift'为'毒药')则更难识别。我们研究多语言大模型(LLMs)在英语-西班牙语、英语-法语、英语-德语场景下对三类词的处理:同形词、非同形词和跨语言同形词。重点评估其在孤立或句法语境中歧义消解与语义判断能力。结果显示,部分模型虽能识别孤立状态下的同形词与非同形词,但在跨语言同形词歧义消解上表现极差,准确率低于随机基准。这表明模型高度依赖拼写相似性而非语义理解。此外,孤立歧义消解性能与真实语义理解无相关性。在语义矛盾句子中,模型对英语与非英语同形词采用不同策略,显示缺乏统一的跨语言歧义处理机制。
原文摘要 · Abstract (English)
Bilingual lexical processing is shaped by the complex interplay of phonological, orthographic, and semantic features of two languages within an integrated mental lexicon. In humans, this is evident in the ease with which cognate words - words similar in both orthographic form and meaning (e.g., blind, meaning "sightless" in both English and German) - are processed, compared to the challenges posed by interlingual homographs, which share orthographic form but differ in meaning (e.g., gift, meaning "present" in English but "poison" in German). We investigate how multilingual Large Language Models (LLMs) handle such phenomena, focusing on English-Spanish, English-French, and English-German cognates, non-cognate, and interlingual homographs. Specifically, we evaluate their ability to disambiguate meanings and make semantic judgments, both when these word types are presented in isolation or within sentence contexts. Our findings reveal that while certain LLMs demonstrate strong performance in recognizing cognates and non-cognates in isolation, they exhibit significant difficulty in disambiguating interlingual homographs, often performing below random baselines. This suggests LLMs tend to rely heavily on orthographic similarities rather than semantic understanding when interpreting interlingual homographs. Further, we find LLMs exhibit difficulty in retrieving word meanings, with performance in isolative disambiguation tasks having no correlation with semantic understanding. Finally, we study how the LLM processes interlingual homographs in incongruent sentences. We find models to opt for different strategies in understanding English and non-English homographs, highlighting a lack of a unified approach to handling cross-lingual ambiguities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。