指令微调让大模型更听用户话,但也更容易接受错误信息。
Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
- 对比基础模型,指令微调后模型更依赖用户输入
- 用户输入错误信息时,微调模型接受率显著升高
- 适合关注模型可靠性与安全性的研究人员
指令微调提升了大语言模型(LLM)遵循用户指令的准确性,改善了可用性并减少了有害输出。然而,这一过程可能增强模型对用户输入的依赖,导致其更易无差别接受错误信息并生成幻觉。现有研究多指出LLM会受外部信息影响而偏离其参数化知识,但鲜有研究直接探讨指令微调对此现象的影响。本研究系统分析了指令微调对LLM misinformation敏感度的影响。结果表明,经过指令微调的LLM在用户提出错误信息时,更倾向于接受并生成相关内容。与基础模型相比,指令微调显著增加了模型对用户输入的依赖,使错误信息的来源从助手角色转移到用户角色。此外,我们还考察了提示结构中用户角色、错误信息长度及系统提示中的警告等因素的影响。研究强调需建立系统性方法,以缓解指令微调带来的意外后果,提升大模型在实际应用中的可靠性。
原文摘要 · Abstract (English)
Instruction-tuning enhances the ability of large language models (LLMs) to follow user instructions more accurately, improving usability while reducing harmful outputs. However, this process may increase the model's dependence on user input, potentially leading to the unfiltered acceptance of misinformation and the generation of hallucinations. Existing studies primarily highlight that LLMs are receptive to external information that contradict their parametric knowledge, but little research has been conducted on the direct impact of instruction-tuning on this phenomenon. In our study, we investigate the impact of instruction-tuning on LLM's susceptibility to misinformation. Our analysis reveals that instruction-tuned LLMs are significantly more likely to accept misinformation when it is presented by the user. A comparison with base models shows that instruction-tuning increases reliance on user-provided information, shifting susceptibility from the assistant role to the user role. Furthermore, we explore additional factors influencing misinformation susceptibility, such as the role of the user in prompt structure, misinformation length, and the presence of warnings in the system prompt. Our findings underscore the need for systematic approaches to mitigate unintended consequences of instruction-tuning and enhance the reliability of LLMs in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。