让大模型推理更可信,通过逻辑规则显式验证结论。
LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning
- 用符号逻辑与大模型结合构建双重上下文,显式处理例外规则。
- 在否定、蕴含和可废止推理上平均提升11.84%,最高增13.2%。
- 适合法律、医疗等需高可信推理的场景,尤其处理复杂例外时。
高可信推理在法律、医疗等关键领域要求结论准确、可验证且明确基于证据。这类推理依赖于从规则、法规和合同中编码的前提,天然涉及可废止或非单调逻辑,单个事实的引入即可推翻普遍规则,带来巨大挑战。尽管大语言模型(LLMs)擅长自然语言处理,但其在标准推理任务中的能力无法满足高可信文本指南所需的严谨推理需求。此类文本中的核心推理常体现否定、蕴含及最关键的可废止规则与例外等特定逻辑结构。本文提出一种新型神经符号架构LOGicalThought(LogT),结合先进逻辑语言与推理器,与LLM协同构建双重符号图上下文与基于逻辑的上下文。该方法将长篇指南的推理问题转化为紧凑的可溯源评估。在四个跨领域基准上对四种基线进行评估,LogT在所有LLM上整体性能提升11.84%。在三种推理模式中表现显著:否定推理提升最多达+10.2%,蕴含推理提升+13.2%,可废止推理提升+5.5%(相较最强基线)。
原文摘要 · Abstract (English)
High-assurance reasoning, particularly in critical domains such as law and medicine, requires conclusions that are accurate, verifiable, and explicitly grounded in evidence. This reasoning relies on premises codified from rules, statutes, and contracts, inherently involving defeasible or non-monotonic logic due to numerous exceptions, where the introduction of a single fact can invalidate general rules, posing significant challenges. While large language models (LLMs) excel at processing natural language, their capabilities in standard inference tasks do not translate to the rigorous reasoning required over high-assurance text guidelines. Core reasoning challenges within such texts often manifest specific logical structures involving negation, implication, and, most critically, defeasible rules and exceptions. In this paper, we propose a novel neurosymbolically-grounded architecture called LOGicalThought (LogT) that uses an advanced logical language and reasoner in conjunction with an LLM to construct a dual symbolic graph context and logic-based context. These two context representations transform the problem from inference over long-form guidelines into a compact grounded evaluation. Evaluated on four multi-domain benchmarks against four baselines, LogT improves overall performance by 11.84% across all LLMs. Performance improves significantly across all three modes of reasoning: by up to +10.2% on negation, +13.2% on implication, and +5.5% on defeasible reasoning compared to the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。