arXiv:2501.14649cs.CL2025-01NAACL

测试大模型在自然语言转正式语言中的拆解与组合能力

Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion

  • 构建DED C框架,分离评估大模型的拆解与组合能力
  • 发现大模型在拆解和组合上均存在明显缺陷
  • 适合关注大模型语言理解与符号系统学习的研究者

为实现通用且鲁棒的自然语言到正式语言转换(N2F),大语言模型需具备在面对陌生正式语言时的强拆解与组合能力,并能应对组合缺口和反直觉符号命名。为探究大模型是否具备这一基础能力,我们提出DED C框架,该框架可半自动完成样本与任务构建,实现对大模型在N2F中拆解与组合能力的解耦评估。基于此框架,我们评估并分析了最先进大模型,主要发现包括:(1)大模型在拆解与组合两方面均表现不足;(2)错误类型广泛,归因于自然语言理解能力弱及符号系统学习与使用能力差;(3)组合缺口与反直觉符号命名均影响大模型的拆解与组合表现。本工作为研究大模型在N2F中的基础能力提供了新视角,缺陷分析与归因有助于后续模型改进。

原文摘要 · Abstract (English)

To achieve generalized and robust natural-to-formal language conversion (N2F), large language models (LLMs) need to have strong capabilities of decomposition and composition in N2F when faced with an unfamiliar formal language and be able to cope with compositional gaps and counter-intuitive symbolic names. To investigate whether LLMs have this set of basic capabilities in N2F, we propose the DEDC framework. This framework semi-automatically performs sample and task construction, allowing decoupled evaluation of the set of decomposition and composition capabilities of LLMs in N2F. Based on this framework, we evaluate and analyze the most advanced LLMs, and the main findings include that: (1) the LLMs are deficient in both decomposition and composition; (2) the LLMs show a wide coverage of error types that can be attributed to deficiencies in natural language understanding and the learning and use of symbolic systems; (3) compositional gaps and counter-intuitive symbolic names both affect the decomposition and composition of the LLMs. Our work provides a new perspective for investigating the basic capabilities of decomposition and composition of LLMs in N2F. The detailed analysis of deficiencies and attributions can help subsequent improvements of LLMs.

自然语言转换大模型能力符号系统语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。