arXiv:2505.23410cs.CL2025-05

发现微调数据不足时,推理阶段提示可弥补事实性差距

From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs

  • 通过系统实验揭示微调数据与推理提示的交互机制
  • 在未知知识上使用提示能显著缩小事实性差距
  • 适合关注模型评估与提示工程的研究者

事实知识提取旨在显式抽取预训练语言模型中参数化的知识以用于下游任务。尽管已有研究探讨监督微调数据对大语言模型事实性的影响,其内在机制仍不清晰。本文通过系统实验重新审视该影响,重点关注在已知与未知知识上微调产生的事实性差距。结果表明,该差距可在推理阶段通过分布外设置或合适的上下文学习提示(如少样本学习和思维链)有效缓解。从知识图谱视角进行理论证明,显示测试阶段提示可能削弱甚至超越微调数据的影响,成为知识提取的主导因素。最终结果揭示了微调数据与测试阶段提示之间的相互作用,表明上下文学习可有效弥补微调数据的不足,并强调需重新思考使用上下文学习提示来评估微调数据选择方法的有效性。

原文摘要 · Abstract (English)

Factual knowledge extraction aims to explicitly extract knowledge parameterized in pre-trained language models for application in downstream tasks. While prior work has been investigating the impact of supervised fine-tuning data on the factuality of large language models (LLMs), its mechanism remains poorly understood. We revisit this impact through systematic experiments, with a particular focus on the factuality gap that arises when fine-tuning on known versus unknown knowledge. Our findings show that this gap can be mitigated at the inference stage, either under out-of-distribution (OOD) settings or by using appropriate in-context learning (ICL) prompts (i.e., few-shot learning and Chain of Thought (CoT)). We prove this phenomenon theoretically from the perspective of knowledge graphs, showing that the test-time prompt may diminish or even overshadow the impact of fine-tuning data and play a dominant role in knowledge extraction. Ultimately, our results shed light on the interaction between finetuning data and test-time prompt, demonstrating that ICL can effectively compensate for shortcomings in fine-tuning data, and highlighting the need to reconsider the use of ICL prompting as a means to evaluate the effectiveness of fine-tuning data selection methods.

大模型提示工程事实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。