用数学本体增强大模型推理,提升专业领域可靠性。
Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge
- 结合数学本体与检索增强生成,注入形式化定义
- 高质量检索下性能提升,低质检索反而降低效果
- 适合需要可验证推理的科研与工程场景
语言模型存在幻觉、脆弱性和缺乏形式化基础等根本缺陷,尤其在高风险专业领域尤为突出。本文以数学为验证场景,构建神经符号系统,利用OpenMath本体通过混合检索与交叉编码重排序,将相关定义注入模型提示。在MATH基准上对三个开源模型的评估显示,当检索质量高时,本体引导的上下文能提升性能;但无关信息会显著降低表现,凸显了神经符号方法的潜力与挑战。
原文摘要 · Abstract (English)
Language models exhibit fundamental limitations -- hallucination, brittleness, and lack of formal grounding -- that are particularly problematic in high-stakes specialist fields requiring verifiable reasoning. I investigate whether formal domain ontologies can enhance language model reliability through retrieval-augmented generation. Using mathematics as proof of concept, I implement a neuro-symbolic pipeline leveraging the OpenMath ontology with hybrid retrieval and cross-encoder reranking to inject relevant definitions into model prompts. Evaluation on the MATH benchmark with three open-source models reveals that ontology-guided context improves performance when retrieval quality is high, but irrelevant context actively degrades it -- highlighting both the promise and challenges of neuro-symbolic approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。