用提问能力评估大模型学新知识的潜力,发现小模型也能很聪明。
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
- 让模型像好奇的人一样对数学公式提问,评估其探索新知的能力。
- 小模型Phi-2在生成问题上表现不逊于大模型GPT-4。
- 提出新评测框架,适合研究AI如何主动学习和提问。
大型语言模型(LLMs)虽存储海量知识,但其获取新知识的潜力尚不明确。本文提出一种新型评估框架,通过引导模型针对引入科学知识的陈述生成问题,模拟初次接触时的好奇心。基于生成问题的质量评分,评估模型的知识获取潜力。我们进行了受控消融实验验证评分方法的有效性,并构建了一个包含1101个物理、化学、数学难度各异的陈述、300个通用知识陈述和567个错误陈述的合成数据集。人类评估验证了模型判断的可靠性,三项指标加权Cohen's kappa约为0.7。结果表明,尽管GPT-4和Mistral 8x7b能生成连贯相关的问题,但较小的Phi-2模型表现相当甚至更优,说明模型大小并非决定知识获取潜力的唯一因素。该框架量化了常被忽视的关键能力,为开发更善于学习的AI系统开辟了新路径。
原文摘要 · Abstract (English)
Large language models (LLMs) can store a massive amount of knowledge, yet their potential to acquire new knowledge remains unknown. We propose a novel evaluation framework that evaluates this capability. This framework prompts LLMs to generate questions about a statement introducing scientific knowledge, simulating a curious person when facing the statement for the first time. We score the qualities of the generated questions, thereby evaluating the knowledge acquisition potential of the LLM. We apply controlled ablation studies to validate our scoring procedures. Additionally, we created a synthetic dataset consisting of 1101 statements in physics, chemistry, and maths with distinct levels of difficulties, 300 general knowledge statements, and 567 incorrect statements. Human evaluations were conducted to validate our model assessments, achieving an approximate weighted Cohen's kappa of 0.7 on all three metrics considered. We find that while large models like GPT-4 and Mistral 8x7b are adept at generating coherent and relevant questions, the smaller Phi-2 model is equally or more effective. This indicates that size does not solely determine a model's knowledge acquisition potential. The proposed framework quantifies a critical model capability that was commonly overlooked and opens up research opportunities for developing more knowledgeable AI systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。