arXiv:2501.00830cs.CLcs.AI2025-01AAAI被引 16

将大模型与动作语言结合,提升复杂行为推理能力。

LLM+AL: Bridging Large Language Models and Action Languages for Complex Reasoning about Actions

  • 用大模型理解语义,动作语言进行符号化推理
  • 在复杂动作推理任务中,准确率显著优于纯大模型
  • 适合需要严谨逻辑推理的智能系统开发

大型语言模型(LLMs)在多种智能任务中取得显著进展,但在需要系统性搜索的复杂行为推理任务上仍表现不足。为此,我们提出一种将大模型自然语言理解能力与动作语言符号推理优势相结合的方法,称为「LLM+AL」。该方法利用大模型在语义解析和常识知识生成方面的优势,同时借助动作语言在编码知识基础上的自动化推理能力。我们在复杂行为推理基准测试中对比了LLM+AL与当前最先进的大模型(包括ChatGPT-4、Claude 3 Opus、Gemini Ultra 1.0和o1-preview)。结果表明,尽管所有方法均存在错误,但经过少量人工修正后,LLM+AL始终能得出正确答案;而独立的大模型即使获得人工反馈也无法改进。此外,该方法还实现了动作语言的自动化生成。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant strides in various intelligent tasks but still struggle with complex action reasoning tasks that require systematic search. To address this limitation, we propose a method that bridges the natural language understanding capabilities of LLMs with the symbolic reasoning strengths of action languages. Our approach, termed "LLM+AL," leverages the LLM's strengths in semantic parsing and commonsense knowledge generation alongside the action language's proficiency in automated reasoning based on encoded knowledge. We compare LLM+AL against state-of-the-art LLMs, including ChatGPT-4, Claude 3 Opus, Gemini Ultra 1.0, and o1-preview, using benchmarks for complex reasoning about actions. Our findings indicate that, although all methods exhibit errors, LLM+AL, with relatively minimal human corrections, consistently leads to correct answers, whereas standalone LLMs fail to improve even with human feedback. LLM+AL also contributes to automated generation of action languages.

大模型动作语言推理符号计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。