揭示日语大模型在多重社会属性交集下的上下文敏感偏见
Intersectional Bias in Japanese Large Language Models from a Contextualized Perspective
- 构建日语交叉偏见评测基准 inter-JBBQ
- 发现相同属性组合下偏见输出随上下文变化
- 适合关注公平性与文化适配的大模型研究者
随着大型语言模型的快速发展,越来越多研究关注其社会偏见问题。尽管多数研究聚焦于单一社会属性的偏见,但社会科学表明,偏见常以交叉性形式出现——即社会属性在特定语境中相互作用产生的偏见。本研究构建了日语交叉偏见评测基准 inter-JBBQ,用于评估大模型在问答任务中的交叉偏见。利用该基准分析 GPT-4o 与 Swallow 模型,发现即使社会属性组合相同,偏见输出仍随上下文动态变化。
原文摘要 · Abstract (English)
An increasing number of studies have examined the social bias of rapidly developed large language models (LLMs). Although most of these studies have focused on bias occurring in a single social attribute, research in social science has shown that social bias often occurs in the form of intersectionality -- the constitutive and contextualized perspective on bias aroused by social attributes. In this study, we construct the Japanese benchmark inter-JBBQ, designed to evaluate the intersectional bias in LLMs on the question-answering setting. Using inter-JBBQ to analyze GPT-4o and Swallow, we find that biased output varies according to its contexts even with the equal combination of social attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。