arXiv:2509.09735cs.CL2025-09被引 1

检测大模型在跨语言决策与摘要中的偏见,提出有效缓解策略。

Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation

  • 构建超大规模多语言提示数据集,测试不同人口属性下的模型表现。
  • 发现模型在决策任务中明显偏袒女性、年轻人和非裔背景者,摘要任务偏见较弱。
  • 新提示指令可降低27%偏见差距,新模型GPT-4o表现更优,适合实际部署前评估。

大型语言模型(LLMs)在各领域的快速应用引发对社会不平等和信息偏见的担忧。本研究聚焦于背景、性别和年龄相关的偏见,考察其在决策与摘要任务中的影响,并分析偏见在英语与荷兰语间的跨语言传播,评估提示引导的缓解策略效果。基于Tamkin等(2023)的数据集并翻译为荷兰语,我们生成了151,200条决策任务提示和176,400条摘要任务提示。在GPT-3.5和GPT-4o上测试了多种人口变量、指令、显著性水平和语言。结果表明,两模型在决策任务中均存在显著偏见,偏好女性、年轻群体及非裔背景;而摘要任务中偏见较弱,但GPT-3.5在英文中表现出显著年龄差异。跨语言分析显示英荷两国偏见模式总体相似,特定类别存在差异。新提出的缓解指令虽未完全消除偏见,但可实现27%的平均偏见差距降低。值得注意的是,与GPT-3.5相反,GPT-4o在所有英文提示中均表现出更低偏见,说明提示缓解在新模型中具有潜力。研究强调需谨慎采纳LLMs,开展上下文特定的偏见检测,亟需持续开发有效缓解策略以确保负责任的AI部署。

原文摘要 · Abstract (English)

The rapid integration of Large Language Models (LLMs) into various domains raises concerns about societal inequalities and information bias. This study examines biases in LLMs related to background, gender, and age, with a focus on their impact on decision-making and summarization tasks. Additionally, the research examines the cross-lingual propagation of these biases and evaluates the effectiveness of prompt-instructed mitigation strategies. Using an adapted version of the dataset by Tamkin et al. (2023) translated into Dutch, we created 151,200 unique prompts for the decision task and 176,400 for the summarisation task. Various demographic variables, instructions, salience levels, and languages were tested on GPT-3.5 and GPT-4o. Our analysis revealed that both models were significantly biased during decision-making, favouring female gender, younger ages, and certain backgrounds such as the African-American background. In contrast, the summarisation task showed minimal evidence of bias, though significant age-related differences emerged for GPT-3.5 in English. Cross-lingual analysis showed that bias patterns were broadly similar between English and Dutch, though notable differences were observed across specific demographic categories. The newly proposed mitigation instructions, while unable to eliminate biases completely, demonstrated potential in reducing them. The most effective instruction achieved a 27\% mean reduction in the gap between the most and least favorable demographics. Notably, contrary to GPT-3.5, GPT-4o displayed reduced biases for all prompts in English, indicating the specific potential for prompt-based mitigation within newer models. This research underscores the importance of cautious adoption of LLMs and context-specific bias testing, highlighting the need for continued development of effective mitigation strategies to ensure responsible deployment of AI.

大模型偏见跨语言决策偏差提示缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。