用大模型生成的推理数据仍含偏见,可能误导后续研究。
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference
- 用GPT-4、Llama-2等生成自然语言推理假设
- 基于假设的分类器准确率达86%-96%,证明存在标注偏见
- 发现大量重复模式,如‘swimming in a pool’超1万次出现于矛盾句
我们检验了用大语言模型替代人工标注者生成自然语言推理(NLI)假设是否导致标注偏差。通过GPT-4、Llama-2和Mistral 7b重制部分斯坦福NLI语料库,并训练仅依赖假设的分类器,以判断生成数据中是否存在标注偏差。在大模型生成的NLI数据集上,基于BERT的假设分类器准确率达到86%-96%,表明这些数据包含显著的假设仅偏见。我们还发现大量“线索”存在于大模型生成的假设中,例如‘swimming in a pool’在GPT-4生成的超过10,000个矛盾句中反复出现。分析证实,已有广泛记录的NLI偏见会持续存在于大模型生成的数据中。
原文摘要 · Abstract (English)
We test whether replacing crowdsource workers with LLMs to write Natural Language Inference (NLI) hypotheses similarly results in annotation artifacts. We recreate a portion of the Stanford NLI corpus using GPT-4, Llama-2 and Mistral 7b, and train hypothesis-only classifiers to determine whether LLM-elicited hypotheses contain annotation artifacts. On our LLM-elicited NLI datasets, BERT-based hypothesis-only classifiers achieve between 86-96% accuracy, indicating these datasets contain hypothesis-only artifacts. We also find frequent "give-aways" in LLM-generated hypotheses, e.g. the phrase "swimming in a pool" appears in more than 10,000 contradictions generated by GPT-4. Our analysis provides empirical evidence that well-attested biases in NLI can persist in LLM-generated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。