arXiv:2510.01219cs.CLcs.AI2025-10

用概念学习任务发现大模型隐含的量词偏向性。

Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset

  • 设计概念学习数据集,通过上下文学习探测模型偏差
  • 发现模型对量词存在向上单调性偏好,且仅在上下文学习中显现
  • 适合关注模型公平性与评估方法的研究者

我们提出一个概念学习任务数据集,用于揭示大语言模型中的隐含偏见。通过上下文概念学习实验,发现语言模型在量词上可能存在向上单调性偏差;该偏差在直接提示测试中不明显,但在引入概念学习组件后显著显现。这表明上下文学习是挖掘模型隐藏偏见的有效手段。

原文摘要 · Abstract (English)

We introduce a dataset of concept learning tasks that helps uncover implicit biases in large language models. Using in-context concept learning experiments, we found that language models may have a bias toward upward monotonicity in quantifiers; such bias is less apparent when the model is tested by direct prompting without concept learning components. This demonstrates that in-context concept learning can be an effective way to discover hidden biases in language models.

大模型偏见概念学习评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。