让大模型用特定语言提问,能答得更准。
Language Specific Knowledge: Do Models Know Better in X than in English?
- 提出'语言专精知识'概念,优化多语言模型的问答语言选择
- 实验证明某些问题在非英语语言下回答准确率更高,如西班牙语问中国相关
- 适合关注多语言模型文化适配性的研究者和开发者
多语言模型通常训练目标是将不同语言中语义相似的内容映射到同一潜在空间。本文揭示了这一目标的细微差别,发现改变输入查询的语言可提升语言模型的问答能力。提出‘语言专精知识’(LSK)概念,指某些问题在特定语言下由模型回答更优,引入语言选择问题:对某些查询,模型在非英语语言下表现更好,甚至在低资源语言中表现更佳。提出多种基线方法(包括自研方法LSKExtractor),在三个包含文化与社会行为规范知识的数据集上评估,结果表明,有策略地选择语言可显著提升模型性能,且最优语言匹配并非直观,例如Gemma模型在西班牙语下对中国的知识掌握最佳,Qwen模型在阿拉伯语和中文下对权威与责任的理解更深入。研究推动开源大模型在部署环境中的文化与语言包容性发展。
原文摘要 · Abstract (English)
Often, multilingual language models are trained with the objective to map semantically similar content (in different languages) in the same latent space. In this paper, we show a nuance in this training objective, and find that by changing the language of the input query, we can improve the question answering ability of language models. We make two main contributions. First, we introduce the term Language Specific Knowledge (LSK) to denote queries that are best answered in an ``expert language'' for a given LLM, thereby enhancing its question-answering ability. We introduce the problem of language selection -- for some queries, language models can perform better when queried in languages other than English, sometimes even better in low-resource languages -- and the goal is to select the optimal language for the query. Second, we introduce a variety of simple to strong baselines to empirically motivate the language selection problem (including one of our own methods called LSKExtractor). During our evaluation, we employ three datasets that contain knowledge about both cultural and social behavioral norms. Overall, the results show that principled language selection can improve the performance of a language model, and that the expected question-to-language map is not always intuitive: Gemma models know most about China and Middle East in Spanish; Qwen models know most about authority and responsibility in Arabic and Chinese. Broadly, our research contributes to the open-source development of language models that are inclusive and more aligned with the cultural and linguistic contexts in which they are deployed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。