测试大模型跨语言语义干扰,发现强模型更晚才用主导语言词义。
Benchmarking Concept-Spilling Across Languages in LLMs
- 通过生成多义词含义,观察模型何时转向主导语言。
- 强模型在序列后期才用主导语言词义,输出更多目标语言真义。
- 适合关注多语言公平性与模型评估的研究者参考。
多语言大模型虽具备出色的跨语言能力,但常对其他语言表征产生系统性偏倚,导致非英语生成时出现语义干扰——我们称之为语言溢出。本文提出一种新型对比评估框架,通过系统测量模型在九种语言中处理高歧义英文词的能力,量化其跨语言语义鲁棒性。方法要求模型准确生成五个意义,结果表明:强模型会更晚依赖主导语言词义,先输出更多目标语言的真实含义;弱模型则早期就转向主导语言。我们使用100个高歧义英语词构建基准测试,评估了多种开源与闭源多语言LLM。结果揭示模型与语言间显著的语义鲁棒性差异,建立了无需归因错误来源的可比排名体系。本研究贡献了一个可扩展的多语言语义评估基准和严谨验证流程,对开发更语言均衡的AI系统至关重要。
原文摘要 · Abstract (English)
Multilingual Large Language Models (LLMs) exhibit remarkable cross-lingual abilities, yet often exhibit a systematic bias toward the representations from other languages, resulting in semantic interference when generating content in non-English languages$-$a phenomenon we define as language spilling. This paper presents a novel comparative framework for evaluating multilingual semantic robustness by systematically measuring how models handle polysemous words across languages. Our methodology provides a relative measure of model performance: when required to generate exactly five meanings, both strong and weak models may resort to meanings from dominant languages, but semantically stronger models do so later in the generation sequence, producing more true meanings from the target language before failing, while weaker models resort to dominant-language meanings earlier in the sequence. We evaluate a diverse set of open and closed multilingual LLMs using a structured meaning generation task across nine languages, employing a carefully curated benchmark of 100 high-polysemy English words. Our findings reveal significant variation in semantic robustness across both models and languages, providing a principled ranking system for model comparison without requiring definitive causal attribution of error sources. We contribute both a scalable comparative benchmark for multilingual semantic evaluation and a rigorous validation pipeline$-$critical tools for developing more linguistically balanced AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。