arXiv:2604.15648cs.CLcs.CV2026-04中稿 · ed

首个评测大模型理解超图能力的基准,推动视觉语言模型在复杂网络推理上的发展。

HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning

论文配图:HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
图 1 · 摘自论文原文
  • 构建首个超图理解与推理评测基准,涵盖84,000个问答样本。
  • 12个先进视觉语言模型在超图任务上表现普遍不足,复杂推理能力受限。
  • 提出通用路由模块WiseHyGR,通过自适应表征提升模型性能,适合图神经网络研究者。

大型视觉语言模型(LVLMs)亟需新领域来拓展其能力边界,但其在超图方面的理解能力仍属空白。现实中,超图在生命科学、社交网络等领域有重要应用。尽管近年来LVLM在复杂拓扑理解上展现潜力,却缺乏评估其在超图上能力的基准,导致其实际边界不明确。为此,本文提出首个超图理解与推理评测基准——HyperGVL,对12个先进LVLM进行系统评估,覆盖84,000个视觉语言问答样本,包含12项任务,从基础组件计数到复杂NP-hard问题推理。数据集包含多尺度合成结构及真实世界引用网络与蛋白质网络。我们还分析了12种文本与视觉超图表示方法,并提出可泛化的路由模块WiseHyGR,通过学习自适应表示显著提升LVLM在超图任务中的表现。本工作为连接超图与大模型开辟了新路径。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) consistently require new arenas to guide their expanding boundaries, yet their capabilities with hypergraphs remain unexplored. In the real world, hypergraphs have significant practical applications in areas such as life sciences and social communities. Recent advancements in LVLMs have shown promise in understanding complex topologies, yet there remains a lack of a benchmark to delineate the capabilities of LVLMs with hypergraphs, leaving the boundaries of their abilities unclear. To fill this gap, in this paper, we introduce $\texttt{HyperGVL}$, the first benchmark to evaluate the proficiency of LVLMs in hypergraph understanding and reasoning. $\texttt{HyperGVL}$ provides a comprehensive assessment of 12 advanced LVLMs across 84,000 vision-language question-answering (QA) samples spanning 12 tasks, ranging from basic component counting to complex NP-hard problem reasoning. The involved hypergraphs contain multiscale synthetic structures and real-world citation and protein networks. Moreover, we examine the effects of 12 textual and visual hypergraph representations and introduce a generalizable router $\texttt{WiseHyGR}$ that improves LVLMs in hypergraph via learning adaptive representations. We believe that this work is a step forward in connecting hypergraphs with LVLMs.

超图视觉语言模型推理评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。