arXiv:2502.17216cs.AIcs.CL2025-02被引 3

选对形式语言,能让大模型推理能力提升显著。

Intermediate Languages Matter: Formal Choice Drives Neurosymbolic LLM Reasoning

  • 用不同形式语言测试推理效果,发现语言选择影响关键表现
  • 上下文编码平均提升推理准确率,注释和标记语法无明显作用
  • 为神经符号推理提供可复现的语言设计标准,适合研究者参考

大型语言模型在众多任务中表现出色,但在形式化推理方面仍显不足。神经符号推理是一种有前景的解决方案:利用大模型将自然语言翻译为形式语言,再由符号求解器生成正确结果。然而,该方法成功的关键因素尚不明确。本文通过在3个数据集上对比6种大模型使用4种形式语言的表现,发现形式语言的选择显著影响语法与语义推理能力。为此,我们提出“中间语言挑战”,即如何为神经符号推理选择合适的正式语言。进一步的消融实验表明,上下文感知的编码有助于大模型推理,而添加注释或使用标记语法则无明显影响。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve astonishing results on a wide range of tasks. However, their formal reasoning ability still lags behind. A promising approach is Neurosymbolic LLM reasoning. It works by using LLMs as translators from natural to formal languages and symbolic solvers for deriving correct results. Still, it remains unclear what the contributing factors to the success of Neurosymbolic LLM reasoning are. This paper shows that one important factor is the choice of the formal language. By comparing 4 formal languages on 3 datasets over 6 LLMs, we show that the choice of formal language affects both the syntactic and the semantic reasoning capability. Thereby, we introduce the intermediate language challenge, which is the challenge of picking a suitable formal language for neurosymbolic reasoning. Further, we compare the effects of using different in-context-learning examples in an ablation study. We conclude that on average, context-aware encodings help LLMs to reason, while there is no apparent effect of using comments or markdown syntax.

大模型推理形式语言神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。