arXiv:2509.24468cs.CL2025-09EMNLP被引 1

用日语数据集发现:去偏方法会严重损害模型文化常识能力

Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset

  • 构建统一评测框架SOBACO,同时评估社会偏见与文化常识
  • 去偏后模型在文化常识任务上最高下降75%准确率
  • 提醒开发者权衡去偏与文化理解,避免过度矫正

大型语言模型存在社会偏见,催生了多种去偏方法。但去偏可能削弱模型能力。以往研究多通过通用语言理解任务评估去偏影响,这些任务与社会偏见关联较弱。而文化常识与社会偏见同根于社会规范与价值观,关联紧密。本文提出SOBACO(SOcial BiAs and Cultural cOmmonsense benchmark)——一个用于统一评测日语模型社会偏见与文化常识的基准。我们在SOBACO上评估多个大模型,发现去偏方法导致模型在文化常识任务上的性能显著下降(最高达75%准确率恶化)。结果表明,需发展兼顾文化常识的去偏方法,以平衡公平性与实用性。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit social biases, prompting the development of various debiasing methods. However, debiasing methods may degrade the capabilities of LLMs. Previous research has evaluated the impact of bias mitigation primarily through tasks measuring general language understanding, which are often unrelated to social biases. In contrast, cultural commonsense is closely related to social biases, as both are rooted in social norms and values. The impact of bias mitigation on cultural commonsense in LLMs has not been well investigated. Considering this gap, we propose SOBACO (SOcial BiAs and Cultural cOmmonsense benchmark), a Japanese benchmark designed to evaluate social biases and cultural commonsense in LLMs in a unified format. We evaluate several LLMs on SOBACO to examine how debiasing methods affect cultural commonsense in LLMs. Our results reveal that the debiasing methods degrade the performance of the LLMs on the cultural commonsense task (up to 75% accuracy deterioration). These results highlight the importance of developing debiasing methods that consider the trade-off with cultural commonsense to improve fairness and utility of LLMs.

去偏文化常识大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。