提出评估大模型推导能力的新框架,发现主流模型推导效果差并提出改进方法。
DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models
- 定义推导关系与推导能力,构建系统评估框架DEVAL。
- GPT-4o等模型在推导应用上表现不佳,平均能力提升15.2%。
- 适合关注模型推理鲁棒性与提示工程的研究者。
评估大语言模型(LLMs)对数据的推理能力仍是开放且紧迫的研究问题。与人类推理不同,人类能基于输入变化推导出输出的相应修改,这种依赖抽象规则的推理模式尚未被充分描述或评估。本文首次正式定义该模式为推导关系(DR),并提出推导能力(DC)概念——即当输入发生特定变化时,能正确修改输出的能力。为此,我们构建了系统性的评估框架DEVAL,用于评测五种主流大模型和一种大型推理模型在七个主流任务中的表现。结果显示,如GPT-4o和Claude3.5等主流模型虽具备中等水平的DR识别能力,但在实际问题求解中应用推导能力显著下降。为进一步提升性能,我们提出新型提示工程方法——推导提示(Derivation Prompting, DP),在所有测试模型上平均提升DC 15.2%,优于现有常用提示技术。
原文摘要 · Abstract (English)
Assessing the reasoning ability of Large Language Models (LLMs) over data remains an open and pressing research question. Compared with LLMs, human reasoning can derive corresponding modifications to the output based on certain kinds of changes to the input. This reasoning pattern, which relies on abstract rules that govern relationships between changes of data, has not been comprehensively described or evaluated in LLMs. In this paper, we formally define this reasoning pattern as the Derivation Relation (DR) and introduce the concept of Derivation Capability (DC), i.e. applying DR by making the corresponding modification to the output whenever the input takes certain changes. To assess DC, a systematically constructed evaluation framework named DEVAL is proposed and used to evaluate five popular LLMs and one Large Reasoning Model in seven mainstream tasks. The evaluation results show that mainstream LLMs, such as GPT-4o and Claude3.5, exhibit moderate DR recognition capabilities but reveal significant drop-offs on applying DR effectively in problem-solving scenarios. To improve this, we propose a novel prompt engineering approach called Derivation Prompting (DP). It achieves an average improvement of 15.2% in DC for all tested LLMs, outperforming commonly used prompt engineering techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。