arXiv:2504.12764cs.LGcs.DM2025-04被引 4

构建首个覆盖多维度的图任务评测框架,揭示LLM在图推理中的表现差异。

GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

  • 设计多类型图、序列化格式与提示策略的综合评测体系
  • Claude-3.5和o4-mini表现最优,但仍有显著提升空间
  • 提出自适应选择机制,适合图结构推理研究者使用

本文提出GraphOmni,一个全面的基准评测框架,用于评估大语言模型在自然语言描述的图论任务中的推理能力。该框架涵盖多种图类型、序列化格式与提示策略,显著超越以往工作在广度与深度上的局限。通过系统性实验,我们发现这些维度之间存在关键交互作用,对模型性能有显著影响。实验表明,SOTA模型如Claude-3.5和o4-mini持续优于其他模型,但仍有明显改进空间。性能波动受具体因素组合影响显著,凸显跨多维综合评估的重要性。此外,序列化与提示策略对开源与闭源模型的影响不同,提示定制化方法的必要性。基于此,我们提出一种受强化学习启发的自适应选择框架,可动态优化影响模型推理的关键因素。该灵活可扩展的基准不仅深化了对LLM在结构化任务中表现的理解,也为未来研究提供坚实基础。代码与数据集已开源:https://github.com/GAI-Community/GraphOmni。

原文摘要 · Abstract (English)

This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni encompasses diverse graph types, serialization formats, and prompting schemes, significantly exceeding prior efforts in both scope and depth. Through extensive systematic evaluation, we identify critical interactions among these dimensions, demonstrating their substantial impact on model performance. Our experiments reveal that state-of-the-art models like Claude-3.5 and o4-mini consistently outperform other models, yet even these leading models exhibit substantial room for improvement. Performance variability is evident depending on the specific combinations of factors we considered, underscoring the necessity of comprehensive evaluations across these interconnected dimensions. Additionally, we observe distinct impacts of serialization and prompting strategies between open-source and closed-source models, encouraging the development of tailored approaches. Motivated by the findings, we also propose a reinforcement learning-inspired framework that adaptively selects the optimal factors influencing LLM reasoning capabilities. This flexible and extendable benchmark not only deepens our understanding of LLM performance on structured tasks but also provides a robust foundation for advancing research in LLM-based graph reasoning. The code and datasets are available at https://github.com/GAI-Community/GraphOmni.

图神经网络大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。