现有图语言模型评估方法失效,新基准揭示其结构推理能力不足。
A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models
- 设计合成图+多模态问题的复合推理基准CLEGR
- 纯语言模型在复杂任务上表现优于带图神经网络的模型
- 揭示当前图语言模型缺乏真正结构化推理能力
图语言模型(GLM)旨在结合图神经网络(GNN)的结构推理能力与大语言模型(LLM)的语义理解能力。然而我们发现,当前主流评估基准主要基于节点分类数据集,无法有效衡量多模态推理能力。分析表明,仅使用单模态信息即可在这些基准上取得优异表现,说明其无需真正实现图-语言融合。为弥补这一评估缺口,我们提出新型基准CLEGR(Compositional Language-Graph Reasoning),通过合成图生成流程与需联合结构与文本语义推理的问题,评估不同复杂度下的多模态推理。对代表性GLM架构的全面评估显示,采用软提示的LLM基线在性能上与完整集成GNN的模型相当。此外,GLMs在需要结构推理的任务中表现显著下降。这些结果表明当前GLMs在图结构推理方面存在明显局限,为推动社区迈向真正的多模态推理提供了基础。
原文摘要 · Abstract (English)
Developments in Graph-Language Models (GLMs) aim to integrate the structural reasoning capabilities of Graph Neural Networks (GNNs) with the semantic understanding of Large Language Models (LLMs). However, we demonstrate that current evaluation benchmarks for GLMs, which are primarily repurposed node-level classification datasets, are insufficient to assess multimodal reasoning. Our analysis reveals that strong performance on these benchmarks is achievable using unimodal information alone, suggesting that they do not necessitate graph-language integration. To address this evaluation gap, we introduce the CLEGR(Compositional Language-Graph Reasoning) benchmark, designed to evaluate multimodal reasoning at various complexity levels. Our benchmark employs a synthetic graph generation pipeline paired with questions that require joint reasoning over structure and textual semantics. We perform a thorough evaluation of representative GLM architectures and find that soft-prompted LLM baselines perform on par with GLMs that incorporate a full GNN backbone. This result calls into question the architectural necessity of incorporating graph structure into LLMs. We further show that GLMs exhibit significant performance degradation in tasks that require structural reasoning. These findings highlight limitations in the graph reasoning capabilities of current GLMs and provide a foundation for advancing the community toward explicit multimodal reasoning involving graph structure and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。