arXiv:2410.21573cs.CLcs.AI2024-10被引 4

用近义词陷阱测试多语大模型,发现其跨语言语义辨析能力仍不足。

Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Sense

  • 以语言相似但含义不同的'伪同源词'设计测试集
  • 多语模型在四组语言对中辨识准确率普遍偏低
  • 揭示高资源语言偏见,适合多语AI公平性研究者参考

多语言大语言模型虽受关注,但在英语之外的可靠性存疑。本文提出新基准StingrayBench,用于评估跨语言词义消歧能力。通过收集印尼语-马来语、印尼语-他加禄语、中文-日语、英语-德语四组语言对中的‘伪同源词’(字形相似但意义迥异的词汇),挑战模型在上下文中正确区分其用法。分析显示,现有模型普遍存在对高资源语言的倾向性偏差。我们还提出了新的量化指标,用于衡量跨语言词义理解与偏差程度。本研究推动更多元包容的语言建模发展,促进多语社区公平使用。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) have gained prominence, but concerns arise regarding their reliability beyond English. This study addresses the gap in cross-lingual semantic evaluation by introducing a novel benchmark for cross-lingual sense disambiguation, StingrayBench. In this paper, we demonstrate using false friends -- words that are orthographically similar but have completely different meanings in two languages -- as a possible approach to pinpoint the limitation of cross-lingual sense disambiguation in LLMs. We collect false friends in four language pairs, namely Indonesian-Malay, Indonesian-Tagalog, Chinese-Japanese, and English-German; and challenge LLMs to distinguish the use of them in context. In our analysis of various models, we observe they tend to be biased toward higher-resource languages. We also propose new metrics for quantifying the cross-lingual sense bias and comprehension based on our benchmark. Our work contributes to developing more diverse and inclusive language modeling, promoting fairer access for the wider multilingual community.

多语言模型词义消歧语言偏见评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。