arXiv:2505.11031cs.CL2025-05被引 7

首个评估大模型符号知识理解与推理能力的基准测试

OntoURL: A Benchmark for Evaluating Large Language Models on Symbolic Ontological Understanding, Reasoning and Learning

  • 构建三维度评测体系,涵盖理解、推理、学习三大能力
  • 覆盖40个本体、8个领域,含57,303道题目,全面检验模型表现
  • 发现模型在理解上优于人类,但在学习任务中仍明显不足

大语言模型在诸多任务中展现出卓越能力,但其处理结构化符号知识的能力仍待深入探索。为此,我们提出本体能力分类体系,并推出首个系统性评估大模型本体理解、推理与学习能力的基准——OntoURL。基于该分类,OntoURL从8个领域中的40个本体中构建了包含57,303个问题的15项任务,对20个开源大模型进行评估。实验显示,当前模型在理解任务上表现良好,但在推理和学习任务中存在明显短板。不同提示策略(少样本与思维链)显著影响性能。人工评估进一步表明,模型在理解和推理上优于人类,但在学习任务中普遍落后。这些发现揭示了大模型在符号知识处理上的潜力与局限,确立了OntoURL作为推动大模型与形式化知识融合的关键基准。

原文摘要 · Abstract (English)

Large language models have demonstrated remarkable capabilities across a wide range of tasks, yet their ability to process structured symbolic knowledge remains underexplored. To address this gap, we propose a taxonomy of ontological capabilities and introduce OntoURL, the first comprehensive benchmark designed to systematically evaluate LLMs' capabilities in handling ontologies -- formal and symbolic representations of domain knowledge. Based on the proposed taxonomy, OntoURL systematically assesses three dimensions: understanding, reasoning, and learning through 15 distinct tasks comprising 57,303 questions derived from 40 ontologies across 8 domains. Experiments with 20 open-source LLMs reveal significant performance differences across models, tasks, and domains, with current LLMs showing capabilities in understanding ontological knowledge but weaknesses in reasoning and learning tasks. Further experiments with few-shot and chain-of-thought prompting illustrate how different prompting strategies affect model performance. Additionally, a human evaluation reveals that LLMs outperform humans in understanding and reasoning tasks but fall short in most learning tasks. These findings highlight both the potential and limitations of LLMs in processing symbolic knowledge and establish OntoURL as a critical benchmark for advancing the integration of LLMs with formal knowledge representations.

大模型评测符号知识本体推理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。