arXiv:2509.13869cs.CL2025-09EMNLP被引 4

大模型在社会偏见判断上未必更准,且偏好特定情境。

Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs

  • 用12个大模型对比不同偏见场景下的价值对齐情况。
  • 参数越大越不靠谱,攻击成功率未随规模下降。
  • 小模型经微调后解释更易读,但一致性较差。

大型语言模型(LLMs)在涉及复杂敏感社会偏见的场景中,可能因与人类价值观不一致而产生不良后果。以往研究多通过专家设计或代理模拟的偏见情景揭示其偏差。然而,不同情境类型(如负面与非负面问题)下,大模型与人类价值观的对齐程度是否存在差异仍不明确。本研究系统考察了12个来自四个模型家族的大模型在四种数据集上的社会偏见价值对齐表现(HVSB)。结果表明,模型参数规模大并不意味着更低的误判率和攻击成功率;不同模型对特定类型情境存在偏好,同一模型家族的判断一致性更高。此外,我们分析了模型对偏见的理解能力及其解释生成效果:各模型对社会偏见的理解无显著差异,且普遍偏好自身生成的解释。我们还为小型语言模型赋予解释能力,微调后的结果表明其生成解释更易读,但模型间共识度较低。

原文摘要 · Abstract (English)

Large language models (LLMs) can lead to undesired consequences when misaligned with human values, especially in scenarios involving complex and sensitive social biases. Previous studies have revealed the misalignment of LLMs with human values using expert-designed or agent-based emulated bias scenarios. However, it remains unclear whether the alignment of LLMs with human values differs across different types of scenarios (e.g., scenarios containing negative vs. non-negative questions). In this study, we investigate the alignment of LLMs with human values regarding social biases (HVSB) in different types of bias scenarios. Through extensive analysis of 12 LLMs from four model families and four datasets, we demonstrate that LLMs with large model parameter scales do not necessarily have lower misalignment rate and attack success rate. Moreover, LLMs show a certain degree of alignment preference for specific types of scenarios and the LLMs from the same model family tend to have higher judgment consistency. In addition, we study the understanding capacity of LLMs with their explanations of HVSB. We find no significant differences in the understanding of HVSB across LLMs. We also find LLMs prefer their own generated explanations. Additionally, we endow smaller language models (LMs) with the ability to explain HVSB. The generation results show that the explanations generated by the fine-tuned smaller LMs are more readable, but have a relatively lower model agreeability.

大模型对齐社会偏见模型解释小模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。