大模型知识易过时不准,新方法提升事实一致性与稳定性
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
- 用实体感知微调构建结构化知识表示,缓解信息不一致
- 24个主流大模型在时间敏感问题上准确率与一致性均不理想
- 适用于需高可靠事实推理的场景,如医疗、法律问答
大语言模型的知识来源于不同时间点和媒介(如维基、社交媒体)的数据快照,这些非结构化信息随时间变化且存在不一致和错误。模型在训练或推理过程中可能因知识漂移导致回答不一致或不准确。本文评估了24个主流大语言模型(包括闭源、部分开源、全开源)作为事实知识库的可靠性,重点考察其在时间敏感问题上的准确性和对提示扰动的鲁棒性。进一步验证了现有方法的效果,并提出一种软神经符号方法——实体感知微调(ENAF),通过在微调阶段引入实体结构化表征,有效降低回答不一致,提升在提示变化下的响应稳定性。
原文摘要 · Abstract (English)
LLMs' sources of knowledge are data snapshots containing factual information about entities collected at different timestamps and from different media types (e.g. wikis, social media, etc.). Such unstructured knowledge is subject to change due to updates through time from past to present. Equally important are the inconsistencies and inaccuracies occurring in different information sources. Consequently, the model's knowledge about an entity may be perturbed while training over the sequence of snapshots or at inference time, resulting in inconsistent and inaccurate model performance. In this work, we study the appropriateness of Large Language Models (LLMs) as repositories of factual knowledge. We consider twenty-four state-of-the-art LLMs that are either closed-, partially (weights), or fully (weight and training data) open-source. We evaluate their reliability in responding to time-sensitive factual questions in terms of accuracy and consistency when prompts are perturbed. We further evaluate the effectiveness of state-of-the-art methods to improve LLMs' accuracy and consistency. We then propose ENtity-Aware Fine-tuning (ENAF), a soft neurosymbolic approach aimed at providing structured representation of entities during fine-tuning to reduce inconsistencies and improve response stability under prompt variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。