构建文本谜题库,评估大模型推理能力边界
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
- 基于Transformer解码器的潜在变量结构设计推理任务
- 提出开源工具enigme生成可定制的文本谜题
- 适合研究模型推理机制或评测系统能力的研究者使用
Transformer解码器语言模型是基于文本生成式人工智能的核心创新。这类模型被广泛部署为多场景通用智能系统,其核心能力在于理解自然语言指令,并利用人类语料库中蕴含的推理知识,对各种新任务执行推理过程。为理解该方法在生成推理方面的局限性,我们提出需关注系统的架构约束。通过对Transformer解码器潜在变量结构的分析,可设计出能探测其推理能力边界的任务。本文提出enigme——一个开源库,用于生成文本类谜题,以训练和评估Transformer解码器模型及未来AI架构的推理能力。
原文摘要 · Abstract (English)
Transformer-decoder language models are a core innovation in text based generative artificial intelligence. These models are being deployed as general-purpose intelligence systems in many applications. Central to their utility is the capacity to understand natural language commands and exploit the reasoning embedded in human text corpora to apply some form of reasoning process to a wide variety of novel tasks. To understand the limitations of this approach to generating reasoning we argue that we need to consider the architectural constraints of these systems. Consideration of the latent variable structure of transformer-decoder models allows us to design reasoning tasks that should probe the boundary of their capacity to reason. We present enigme, an open-source library for generating text-based puzzles to be used in training and evaluating reasoning skills within transformer-decoder models and future AI architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。