arXiv:2502.17945cs.CL2025-02中稿 · ACL

首次系统评估大模型在多语言场景下的国家偏见,发现英语主导仍普遍存在。

Assessing Agentic Large Language Models in Multilingual National Bias

  • 通过大学申请、旅行、搬迁三类任务测试多语言推理偏差
  • GPT-4与Sonnet在英语国家表现更好,但未实现真正多语言对齐
  • 链式思维提示能缓解部分偏见,适合跨文化应用研究者参考

大语言模型在多语言自然语言处理中备受关注,但关于跨语言偏见的研究仍局限于即时上下文偏好。基于推理的推荐在不同语言间的差异尚未被充分探索,甚至缺乏描述性分析。本研究首次填补这一空白,通过大学申请、旅行和搬迁三个关键场景,测试LLM在跨语言个性化建议中的适用性与能力。我们分析了先进LLM在多语言决策任务中的表现,量化其生成评分中的偏见,并评估人口因素与推理策略(如链式思维提示)对偏见模式的影响。结果表明,本地语言偏见在各类任务中普遍存在;尽管GPT-4和Sonnet相较于GPT-3.5在英语国家降低了偏见,但未能实现稳健的多语言对齐,凸显了多语言AI代理在教育等领域的广泛影响。代码已开源。

原文摘要 · Abstract (English)

Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, while studies on risks associated with cross biases are limited to immediate context preferences. Cross-language disparities in reasoning-based recommendations remain largely unexplored, with a lack of even descriptive analysis. This study is the first to address this gap. We test LLM's applicability and capability in providing personalized advice across three key scenarios: university applications, travel, and relocation. We investigate multilingual bias in state-of-the-art LLMs by analyzing their responses to decision-making tasks across multiple languages. We quantify bias in model-generated scores and assess the impact of demographic factors and reasoning strategies (e.g., Chain-of-Thought prompting) on bias patterns. Our findings reveal that local language bias is prevalent across different tasks, with GPT-4 and Sonnet reducing bias for English-speaking countries compared to GPT-3.5 but failing to achieve robust multilingual alignment, highlighting broader implications for multilingual AI agents and applications such as education. \footnote{Code available at: https://github.com/yiyunya/assess_agentic_national_bias

多语言偏见大模型评估推理偏差AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。