arXiv:2501.11968cs.AIcs.LG2025-01被引 7

用图像化图结构让大模型像人一样解复杂组合问题

Bridging Visualization and Optimization: Multimodal Large Language Models on Graph-Structured Combinatorial Optimization

  • 将图结构转为图像,利用多模态大模型的空间推理能力
  • 在6类图任务上表现超越传统方法,无需复杂训练
  • 适合想用简单方式解决图优化难题的研究者

图结构组合优化问题因其非线性和复杂性,传统计算方法常失效或成本过高。本文提出将图转化为图像以保留其高阶结构特征,使机器能模仿人类的空间推理能力。结合多模态大语言模型(MLLMs)与简单搜索策略,构建新框架应对多种图任务,涵盖影响最大化、网络瓦解等序列决策问题及六类基础图问题。实验表明,MLLMs展现出卓越空间智能,能以类人直觉高效处理复杂图数据。结果表明,将MLLMs与简单优化策略结合,可在无复杂推导、无需高算力训练的情况下,有效求解图结构组合问题。

原文摘要 · Abstract (English)

Graph-structured combinatorial challenges are inherently difficult due to their nonlinear and intricate nature, often rendering traditional computational methods ineffective or expensive. However, these challenges can be more naturally tackled by humans through visual representations that harness our innate ability for spatial reasoning. In this study, we propose transforming graphs into images to preserve their higher-order structural features accurately, revolutionizing the representation used in solving graph-structured combinatorial tasks. This approach allows machines to emulate human-like processing in addressing complex combinatorial challenges. By combining the innovative paradigm powered by multimodal large language models (MLLMs) with simple search techniques, we aim to develop a novel and effective framework for tackling such problems. Our investigation into MLLMs spanned a variety of graph-based tasks, from combinatorial problems like influence maximization to sequential decision-making in network dismantling, as well as addressing six fundamental graph-related issues. Our findings demonstrate that MLLMs exhibit exceptional spatial intelligence and a distinctive capability for handling these problems, significantly advancing the potential for machines to comprehend and analyze graph-structured data with a depth and intuition akin to human cognition. These results also imply that integrating MLLMs with simple optimization strategies could form a novel and efficient approach for navigating graph-structured combinatorial challenges without complex derivations, computationally demanding training and fine-tuning.

图神经网络多模态模型组合优化空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。