大模型更抗误导,能更好结合提示与内部知识
Too Big to Fool: Resisting Deception in Language Models
- 比较同系列不同规模模型对误导性提示的响应
- 大模型在抗欺骗和执行指令上均表现更优
- 优势源于对提示隐含信息的更好利用
大型语言模型需在权重编码的知识与提示中的上下文信息之间取得平衡,以生成准确回答。本文通过分析同一模型家族中不同容量的模型如何处理故意误导的上下文信息,研究这一平衡机制。实验表明,更大的模型对误导性提示具有更强的抵抗力,展现出更先进的提示信息解读与整合能力。此外,大模型在遵循合法指令方面也优于小模型,说明其抗干扰能力并非因忽略上下文信息所致。该现象很可能并非源于记忆,而是源于模型能更有效地利用提示中隐含的任务相关信息与其内部存储知识相结合。
原文摘要 · Abstract (English)
Large language models must balance their weight-encoded knowledge with in-context information from prompts to generate accurate responses. This paper investigates this interplay by analyzing how models of varying capacities within the same family handle intentionally misleading in-context information. Our experiments demonstrate that larger models exhibit higher resilience to deceptive prompts, showcasing an advanced ability to interpret and integrate prompt information with their internal knowledge. Furthermore, we find that larger models outperform smaller ones in following legitimate instructions, indicating that their resilience is not due to disregarding in-context information. We also show that this phenomenon is likely not a result of memorization but stems from the models' ability to better leverage implicit task-relevant information from the prompt alongside their internally stored knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。