arXiv:2604.00568cs.CL2026-04

构建日语社会偏见评估数据集,聚焦推理中的归因偏差。

A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory

  • 基于归因理论设计新数据集,固定结论考察推理过程偏见
  • 含216个日本文化特有案例,可更敏感检测模型差异
  • 适合研究日语模型公平性与文化适配的学者使用

提升大语言模型公平性时,评估特定语言地区文化背景下的社会偏见至关重要。然而,现有日语基准多依赖英文数据翻译,未必适配日本文化;且仅评估结论阶段的偏见,忽略推理过程中的隐性偏见。本研究基于社会心理学中的归因理论,构建新数据集 JUBAKU-v2,通过固定结论,评估模型在归因行为时对内群体与外群体的偏见。该数据集包含216个反映日本文化特征的示例。实验表明,相比现有基准,JUBAKU-v2能更敏感地检测不同模型间的性能差异。

原文摘要 · Abstract (English)

In enhancing the fairness of Large Language Models (LLMs), evaluating social biases rooted in the cultural contexts of specific linguistic regions is essential. However, most existing Japanese benchmarks heavily rely on translating English data, which does not necessarily provide an evaluation suitable for Japanese culture. Furthermore, they only evaluate bias in the conclusion, failing to capture biases lurking in the reasoning. In this study, based on attribution theory in social psychology, we constructed a new dataset, ``JUBAKU-v2,'' which evaluates the bias in attributing behaviors to in-groups and out-groups within reasoning while fixing the conclusion. This dataset consists of 216 examples reflecting cultural biases specific to Japan. Experimental results verified that it can detect performance differences across models more sensitively than existing benchmarks.

社会偏见日语模型归因理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。