arXiv:2410.05298cs.LGcs.AI2024-10被引 33

测试大模型理解图结构模式的能力,发现其表现依赖输入格式。

How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension

  • 设计新基准评估大模型对图模式的理解能力
  • O1-mini在多数任务中表现最佳,输入格式影响性能
  • 大模型的推理方式与传统算法不同,适合图模式探索

评测大语言模型(LLMs)在图相关任务中的能力与局限性正成为研究热点。现有研究表明,LLMs初步具备理解图结构和节点特征的能力,但在图模式挖掘方面的潜力仍待探索,而这是计算化学、生物学与社交网络分析等领域的关键环节。为此,本文提出一个全面的基准,用于评估LLMs在基于术语或拓扑描述的图模式理解能力,以及从数据中自主发现图模式的能力。该基准涵盖合成与真实数据集,涉及11项任务和7种模型,实验框架支持扩展。结果表明:(1)LLMs具备初步的图模式理解能力,其中O1-mini在多数任务中表现最优;(2)将输入数据格式化以匹配预训练知识可提升性能;(3)LLMs采用的策略可能不同于传统算法。

原文摘要 · Abstract (English)

Benchmarking the capabilities and limitations of large language models (LLMs) in graph-related tasks is becoming an increasingly popular and crucial area of research. Recent studies have shown that LLMs exhibit a preliminary ability to understand graph structures and node features. However, the potential of LLMs in graph pattern mining remains largely unexplored. This is a key component in fields such as computational chemistry, biology, and social network analysis. To bridge this gap, this work introduces a comprehensive benchmark to assess LLMs' capabilities in graph pattern tasks. We have developed a benchmark that evaluates whether LLMs can understand graph patterns based on either terminological or topological descriptions. Additionally, our benchmark tests the LLMs' capacity to autonomously discover graph patterns from data. The benchmark encompasses both synthetic and real datasets, and a variety of models, with a total of 11 tasks and 7 models. Our experimental framework is designed for easy expansion to accommodate new models and datasets. Our findings reveal that: (1) LLMs have preliminary abilities to understand graph patterns, with O1-mini outperforming in the majority of tasks; (2) Formatting input data to align with the knowledge acquired during pretraining can enhance performance; (3) The strategies employed by LLMs may differ from those used in conventional algorithms.

图神经网络大模型评估模式识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。