自一致性提升百科知识召回,且在GPT-4o上达到89%准确率
Does Self-Consistency Improve the Recall of Encyclopedic Knowledge?

- 用数据驱动方法划分MMLU为推理与知识回忆子集
- 自一致性在两类任务中均提升性能,知识召回达89%
- 适合关注大模型知识能力评估的研究者
尽管自一致性已知可提升符号推理表现,但其对百科知识召回的影响尚不明确,因缺乏针对性评估基准。为此,我们基于先前工作中的数据驱动启发法,在主流的MMLU基准上构建了知识召回子集。通过验证发现,该子集在符号推理与知识回忆上的性能模式分别与GSM8K和MedMCQA一致。在此可靠基础上,我们发现自一致性在两类任务中均有持续提升效果,尽管其基于思维链(CoT)提示主要针对符号推理。最终,我们使用GPT-4o在MMLU上实现了89%的准确率,创下当前最佳纪录。
原文摘要 · Abstract (English)
While self-consistency is known to improve performance on symbolic reasoning, its effect on the recall of encyclopedic knowledge is unclear due to a lack of targeted evaluation grounds. To address this, we establish such a knowledge recall split for the popular MMLU benchmark by applying a data-driven heuristic from prior work. We validate this split by showing that the performance patterns on the symbolic reasoning and knowledge recall subsets mirror those of GSM8K and MedMCQA, respectively. Using this solid ground, we find that self-consistency consistently improves performance across both symbolic reasoning and knowledge recall, even though its underlying CoT prompting is primarily effective for symbolic reasoning. As a result, we achieve an 89\% accuracy on MMLU, the best performance to date with the use of GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。