用神经符号方法提升逻辑推理模型在扰动下的鲁棒性
Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- 将大模型与符号逻辑求解器结合,增强推理一致性
- 神经符号方法在对抗性变体上表现更稳定,但整体准确率较低
- 结合思维链提示后仍不如纯神经模型,提示新方向
当前自然语言处理研究致力于提升大语言模型(LLMs)的推理能力,尤其关注其在任务变化下的泛化与鲁棒性。反事实任务变体通过微小但语义有意义的修改(如更改单个谓词或交换常量角色),对有效的一阶逻辑(FOL)问题实例进行扰动,以检验推理系统在扰动下保持逻辑一致性的能力。先前研究表明,LLMs在反事实变体上变得脆弱,表明其常依赖表面模式生成答案。本文探讨神经符号(NS)方法——将LLM与符号逻辑求解器结合——是否能缓解此问题。在不同规模的LLM上实验发现,NS方法更具鲁棒性,但整体性能低于纯神经方法。随后提出NSCoT,融合NS方法与思维链(CoT)提示,虽提升了性能,但仍落后于标准CoT。分析为未来研究提供了新方向。
原文摘要 · Abstract (English)
Recent trends in NLP aim to improve reasoning capabilities in Large Language Models (LLMs), with key focus on generalization and robustness to variations in tasks. Counterfactual task variants introduce minimal but semantically meaningful changes to otherwise valid first-order logic (FOL) problem instances altering a single predicate or swapping roles of constants to probe whether a reasoning system can maintain logical consistency under perturbation. Previous studies showed that LLMs becomes brittle on counterfactual variations, suggesting that they often rely on spurious surface patterns to generate responses. In this work, we explore if a neurosymbolic (NS) approach that integrates an LLM and a symbolic logical solver could mitigate this problem. Experiments across LLMs of varying sizes show that NS methods are more robust but perform worse overall that purely neural methods. We then propose NSCoT that combines an NS method and Chain-of-Thought (CoT) prompting and demonstrate that while it improves performance, NSCoT still lags behind standard CoT. Our analysis opens research directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。