arXiv:2510.19616cs.CL2025-10被引 1

构建首个波斯语社会偏见评估数据集,助力大模型文化适配

PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models

  • 通过人机协作设计16类文化议题问卷,覆盖37000+问题
  • 发现现有大模型在波斯语场景下存在显著社会偏见
  • 适合研究多语言公平性、文化适应性模型的学者使用

随着大语言模型(LLMs)广泛应用,确保其与社会规范对齐成为关键挑战。尽管已有研究关注多种语言中的偏见检测,但针对波斯文化背景的社会偏见资源仍严重不足。本文提出PBBQ,一个旨在评估波斯语LLMs社会偏见的综合性基准数据集。该数据集涵盖16个文化类别,基于250名不同背景人员完成的问卷,并与社会科学专家紧密合作以确保有效性。最终生成的PBBQ包含超过37,000条精心筛选的问题,为波斯语模型的偏见评估与缓解提供基础。我们在PBBQ上测试了多个开源模型、一个闭源模型及特定于波斯语的微调模型。结果表明,当前模型在波斯文化背景下普遍存在显著社会偏见。此外,通过对比模型输出与人类回答,我们发现模型常复现人类偏见模式,揭示了学习表征与文化刻板印象之间的复杂互动。论文被接受后,PBBQ数据集将公开供后续研究使用。内容警告:本文包含不当内容。

原文摘要 · Abstract (English)

With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap in resources addressing social biases within Persian cultural contexts. In this work, we introduce PBBQ, a comprehensive benchmark dataset designed to evaluate social biases in Persian LLMs. Our benchmark, which encompasses 16 cultural categories, was developed through questionnaires completed by 250 diverse individuals across multiple demographics, in close collaboration with social science experts to ensure its validity. The resulting PBBQ dataset contains over 37,000 carefully curated questions, providing a foundation for the evaluation and mitigation of bias in Persian language models. We benchmark several open-source LLMs, a closed-source model, and Persian-specific fine-tuned models on PBBQ. Our findings reveal that current LLMs exhibit significant social biases across Persian culture. Additionally, by comparing model outputs to human responses, we observe that LLMs often replicate human bias patterns, highlighting the complex interplay between learned representations and cultural stereotypes.Upon acceptance of the paper, our PBBQ dataset will be publicly available for use in future work. Content warning: This paper contains unsafe content.

偏见评估多语言波斯语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。