小模型用大模型生成的数据微调后幻觉增多,验证了知识不匹配假说。
Exploring the Knowledge Mismatch Hypothesis: Hallucination Propensity in Small Models Fine-tuned on Data from Larger Models
- 用大模型生成的数据微调小模型,导致知识输入与模型已有知识不匹配。
- 在未见测试集上,该方法使错误回答率显著高于小模型自生成数据微调的结果。
- 研究揭示了模型幻觉的潜在机制,适合关注AI安全与可信性的研究者参考。
近期大量小规模语言模型通过使用大模型生成的数据进行微调而诞生,这些小模型能产生与大模型相似的输出质量。然而,其主要缺陷之一是比大模型更容易产生幻觉,即生成看似连贯但事实错误的信息,传播虚假、有毒内容和刻板印象。其中一种可能原因是:用大模型生成的数据微调小模型会引发知识不匹配,即输入数据的知识与小模型自身知识图谱之间存在错位。本文通过实验发现,在未见过的测试集上,使用大模型数据微调的小模型产生的错误答案明显多于使用小模型自身生成数据微调的版本,验证了知识不匹配假说。
原文摘要 · Abstract (English)
Recently, there has been an explosion of large language models created through fine-tuning with data from larger models. These small models able to produce outputs that appear qualitatively similar to significantly larger models. However, one of the key limitations that have been observed with these models is their propensity to hallucinate significantly more often than larger models. In particular, they have been observed to generate coherent outputs that involve factually incorrect information and spread misinformation, toxicity, and stereotypes. There are many potential causes of hallucination, of which, one hypothesis is that fine-tuning a model on data produced by a larger model leads to a knowledge mismatch which contributes to hallucination. In particular, it is hypothesized that there is a mismatch between the knowledge that is fed to the model to fine-tune it and the knowledge that is already present in the graph. Fine-tuning the model on data that has such mismatch could contribute to an increased propensity to hallucinate. We show that on an unseen test set, a smaller model fine-tuned on data generated from a larger model produced more wrong answers when compared to models fine-tuned on data created by the small model, which confirms the hypothesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。