低资源语言查询时大模型回答质量下降,且语言影响文化背景选择。
Language Models Entangle Language and Culture
- 基于真实对话数据构建多语言开放问题集进行评估
- 低资源语言答案质量显著低于高资源语言(如英语)
- 语言选择会改变模型使用的文化背景,影响回答质量
用户不应因使用不同语言而系统性处于不利地位;即无论使用何种语言,交互质量应保持一致。本文基于对WildChat数据集的分析,构建了一组真实世界的开放式问题,用于评估不同语言下模型回答质量是否存在差异,特别是回答质量是否依赖于查询语言。同时,通过LLM-as-a-Judge方法识别响应中的文化语境,探究语言与文化在大模型中的纠缠现象。进一步在多语言翻译版CulturalBench基准上评估模型表现。结果表明,大模型在低资源语言上的回答质量持续偏低,且语言选择显著影响模型所采用的文化背景,进而影响下游回答质量。
原文摘要 · Abstract (English)
Users should not be systemically disadvantaged by the language they use for interacting with LLMs; i.e. users across languages should get responses of similar quality irrespective of language used. In this work, we create a set of real-world open-ended questions based on our analysis of the WildChat dataset and use it to evaluate whether responses vary by language, specifically, whether answer quality depends on the language used to query the model. We also investigate how language and culture are entangled in LLMs such that choice of language changes the cultural information and context used in the response by using LLM-as-a-Judge to identify the cultural context present in responses. To further investigate this, we evaluate LLMs on a translated subset of the CulturalBench benchmark across multiple languages. Our evaluations reveal that LLMs consistently provide lower quality answers to open-ended questions in low resource languages. We find that language significantly impacts the cultural context used by the model. This difference in context impacts the quality of the downstream answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。