提升大模型常识推理能力,通过筛选噪声知识和修正错误推理链。
LINKED: Eliciting, Filtering and Integrating Knowledge in Large Language Model for Commonsense Reasoning
- 设计奖励模型过滤噪声知识,用一致性模块减少无效推理。
- 在两个基准上准确率提升最高达9.0%,优于当前最优方法。
- 提出新评估指标衡量知识注入的正负影响,适合研究常识推理者。
大语言模型在知识密集型任务中表现不佳,常识推理即为典型问题。现有方法多通过知识图谱检索或自增强机制激发模型内部知识,但噪声知识与无效推理仍导致答案不准。为此,我们提出新型方法LINKED,通过奖励模型过滤噪声知识,并引入边际一致推理模块减少无效推理。在两个复杂常识推理基准上的综合实验表明,该方法准确率最高提升9.0%,超越当前最优基线。此外,为评估知识注入的正负影响,我们提出有效性保持分数(effectiveness-preservation score)作为新度量标准。通过大量实验,我们深入分析了大模型在常识推理中的行为模式,获得多项有意义结论。
原文摘要 · Abstract (English)
Large language models (LLMs) sometimes demonstrate poor performance on knowledge-intensive tasks, commonsense reasoning is one of them. Researchers typically address these issues by retrieving related knowledge from knowledge graphs or employing self-enhancement methods to elicit knowledge in LLMs. However, noisy knowledge and invalid reasoning issues hamper their ability to answer questions accurately. To this end, we propose a novel method named eliciting, filtering and integrating knowledge in large language model (LINKED). In it, we design a reward model to filter out the noisy knowledge and take the marginal consistent reasoning module to reduce invalid reasoning. With our comprehensive experiments on two complex commonsense reasoning benchmarks, our method outperforms SOTA baselines (up to 9.0% improvement of accuracy). Besides, to measure the positive and negative impact of the injected knowledge, we propose a new metric called effectiveness-preservation score for the knowledge enhancement works. Finally, through extensive experiments, we conduct an in-depth analysis and find many meaningful conclusions about LLMs in commonsense reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。