用概念学习任务发现大模型隐含的量词偏向性。
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
- 设计概念学习数据集,通过上下文学习探测模型偏差
- 发现模型对量词存在向上单调性偏好,且仅在上下文学习中显现
- 适合关注模型公平性与评估方法的研究者
我们提出一个概念学习任务数据集,用于揭示大语言模型中的隐含偏见。通过上下文概念学习实验,发现语言模型在量词上可能存在向上单调性偏差;该偏差在直接提示测试中不明显,但在引入概念学习组件后显著显现。这表明上下文学习是挖掘模型隐藏偏见的有效手段。
原文摘要 · Abstract (English)
We introduce a dataset of concept learning tasks that helps uncover implicit biases in large language models. Using in-context concept learning experiments, we found that language models may have a bias toward upward monotonicity in quantifiers; such bias is less apparent when the model is tested by direct prompting without concept learning components. This demonstrates that in-context concept learning can be an effective way to discover hidden biases in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。