揭示全局注意力图模型在混合整数规划中的表达能力边界
A Weisfeiler-Leman Characterization of Global-Attention Graph Transformers for Mixed-Integer Linear Programs

- 用1维威斯费勒-莱曼测试分析模型表达力,发现其受制于该测试
- 十种主流图编码器在1-WL等价图上均生成相同嵌入
- 适合关注模型表达局限性的优化与机器学习研究者
具有全局注意力的图基础模型(GFMs)被广泛用于表示混合整数线性规划(MILPs),旨在捕捉标准图神经网络无法建模的全局结构。本文通过图同构测试研究其表达能力,探究哪些MILP实例会被映射为相同的表示。证明了一类结合全局线性注意力、边权重交叉注意力和二分消息传递的分层图变压器,其表达力受限于一维威斯费勒-莱曼(1-WL)测试:在任意参数设置下,1-WL等价的MILP图会获得相同图嵌入。组合式证明表明,每个架构组件均为对称多重集函数,因而保持1-WL等价性。在十种不同图编码器中验证了这一特性,包括Graphormer、GraphGPS、Set-Transformer和Gasse风格模型。无论模型容量、图规模或池化操作如何,所有测试编码器均将1-WL等价但非同构的图对映射为数值相同的嵌入。因此,1-WL等价类内部变化的图不变量无法从这些表示中恢复。进一步表明,超越1-WL的表达力源于输入编码而非注意力机制:随机游走位置编码可区分构造出的图对,而其他构造揭示了该方法的局限性。本研究刻画了全局注意力图基础模型的表达力,并提供了一种无须依赖具体编码器的诊断方法以检测1-WL引发的表示等价。
原文摘要 · Abstract (English)
Graph foundation models (GFMs) with global attention are increasingly used to represent mixed-integer linear programs (MILPs), aiming to capture structure beyond the locality of standard graph neural networks. We study their expressive power through graph isomorphism testing, asking which MILP instances they map to identical representations. We prove that a broad class of hierarchical graph transformers combining global linear attention, edge-weighted cross-attention, and bipartite message passing is bounded by the one-dimensional Weisfeiler-Leman (1-WL) test: under any parameter setting, 1-WL-equivalent MILP graphs receive identical graph embeddings. Our compositional proof shows that each architectural component is a symmetric multiset function and thus preserves 1-WL equivalence. We validate this characterization across ten diverse graph encoders, including Graphormer-, GraphGPS-, Set-Transformer-, and Gasse-style models. Across model capacities, graph scales, and pooling operators, every tested encoder maps 1-WL-equivalent non-isomorphic graph pairs to numerically identical embeddings. Consequently, graph invariants that vary within a 1-WL equivalence class cannot be recovered from these representations. We further show that expressiveness beyond 1-WL arises from input encoding rather than attention: random-walk positional encodings separate the constructed pairs, while additional constructions expose the limits of this remedy. These results characterize the expressive power of global-attention GFMs and provide an encoder-agnostic diagnostic for detecting 1-WL-induced representation equivalence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。