新评测方法揭示大模型真实逻辑理解能力,对话型模型表现优异。
Do LLMs Really Struggle at NL-FOL Translation? Revealing their Strengths via a Novel Benchmarking Strategy
- 设计新评测协议,区分深层理解与表面模式匹配
- 顶尖对话型大模型在句级逻辑翻译中表现良好
- 适合关注逻辑推理与形式化验证的研究者
一阶逻辑(FOL)因其表达力强且语义明确,是将自然语言(NL)概念形式化的重要工具,可用于系统属性的规范与验证。尽管将FOL转为可读英语相对简单,但反向的自然语言到一阶逻辑(NL-FOL)翻译对人类和机器而言仍是长期挑战。尽管大语言模型(LLMs)的出现带来了突破希望,但现有研究对其能力的评估结果矛盾。本文提出三重贡献:首先,批判性分析现有数据集与评测协议,揭示其可能误导对LLMs真实能力的判断;其次,提出一种新型评测协议,专门用于区分真正的语义逻辑理解与表层模式识别、记忆或数据污染;第三,通过新方法证明,当前最先进的对话导向型大模型展现出强大的NL-FOL翻译能力及对句级逻辑的真实掌握,而以嵌入为中心的模型表现显著更差。
原文摘要 · Abstract (English)
Due to its expressiveness and unambiguous nature, First-Order Logic (FOL) is a powerful formalism for representing concepts expressed in natural language (NL). This is useful, e.g., for specifying and verifying desired system properties. While translating FOL into human-readable English is relatively straightforward, the inverse problem, converting NL to FOL (NL-FOL translation), has remained a longstanding challenge, for both humans and machines. Although the emergence of Large Language Models (LLMs) promised a breakthrough, recent literature provides contrasting results on their ability to perform NL-FOL translation. In this work, we provide a threefold contribution. First, we critically examine existing datasets and protocols for evaluating NL-FOL translation performance, revealing key limitations that may cause a misrepresentation of LLMs' actual capabilities. Second, to overcome these shortcomings, we propose a novel evaluation protocol explicitly designed to distinguish genuine semantic-level logical understanding from superficial pattern recognition, memorization, and dataset contamination. Third, using this new approach, we show that state-of-the-art, dialogue-oriented LLMs demonstrate strong NL-FOL translation skills and a genuine grasp of sentence-level logic, whereas embedding-centric models perform markedly worse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。