用真实问卷数据增强上下文,让大模型政治偏见测试更稳定可靠。
Leveraging In-Context Learning for Political Bias Testing of LLMs
- 用人类调查数据作上下文示例,提升偏见测试稳定性。
- 发现指令微调会改变模型偏见方向,且大模型偏见更小。
- 适合关注模型公平性与评测方法改进的研究者使用。
大量研究通过向大模型提问来评估其潜在政治偏见,但此类探测方法稳定性差,导致模型间比较不可靠。本文认为大模型需要更多上下文信息,提出新探测任务Questionnaire Modeling(QM),利用真实人类调查数据作为上下文示例。实验表明,QM显著提升了基于问题的偏见评估稳定性,并可用于比较指令微调模型与其基础版本的偏见差异。对不同规模模型的测试显示,指令微调确实可能改变偏见方向;同时观察到,更大模型能更有效利用上下文示例,在QM中普遍表现出更低的偏见分数。数据与代码已公开。
原文摘要 · Abstract (English)
A growing body of work has been querying LLMs with political questions to evaluate their potential biases. However, this probing method has limited stability, making comparisons between models unreliable. In this paper, we argue that LLMs need more context. We propose a new probing task, Questionnaire Modeling (QM), that uses human survey data as in-context examples. We show that QM improves the stability of question-based bias evaluation, and demonstrate that it may be used to compare instruction-tuned models to their base versions. Experiments with LLMs of various sizes indicate that instruction tuning can indeed change the direction of bias. Furthermore, we observe a trend that larger models are able to leverage in-context examples more effectively, and generally exhibit smaller bias scores in QM. Data and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。