动态选择图结构表示,提升视觉语言模型的图问答能力。
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
- 根据查询动态选择最优图拓扑表示方式
- 在无需额外训练下提升算法问答与真实任务表现
- 适用于多种模型、任务和领域,灵活可扩展
视觉语言模型(VLMs)已成为跨领域零样本问答的通用解决方案。然而,使VLMs有效理解结构化图并进行准确高效的图问答仍具挑战性。现有方法通常依赖单一图拓扑表示(GTR),如固定风格的图像或统一文本描述,这种“一刀切”策略常忽略模型和任务特异性偏好,导致回答不准确或过长。为此,我们提出DynamicGTR框架,在推理时动态为每个查询选择最优的GTR,从而在可定制的准确率与简洁性权衡下增强VLM的零样本图问答能力。大量实验表明,DynamicGTR不仅提升了基于VLM的图算法问答性能,还能将合成图算法任务中学到的经验迁移至真实应用(如链接预测与节点分类),且无需额外训练。此外,该方法在任务、领域和模型间展现出强泛化能力,具备作为广泛图场景通用解决方案的潜力。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have emerged as versatile solutions for zero-shot question answering (QA) across various domains. However, enabling VLMs to effectively comprehend structured graphs and perform accurate, efficient QA remains challenging. Existing approaches typically rely on one single graph topology representation (GTR), such as fixed-style visual images or unified text descriptions. This ``one-size-fits-all'' strategy often neglects model-specific and task-specific preferences, resulting in inaccurate or over-lengthy responses to graph-related queries. To address this, we propose the $\mbox{DynamicGTR}$ framework, which dynamically selects the optimal GTR for each query during inference, thereby enhancing the zero-shot graph QA capabilities of VLMs with a customizable accuracy and brevity trade-off. Extensive experiments show that DynamicGTR not only improves VLM-based graph algorithm QA performance but also successfully transfers the experience trained from synthetic graph algorithm tasks to real-world applications like link prediction and node classification, without any additional training. Additionally, DynamicGTR demonstrates strong transferability across tasks, domains, and models, suggesting its potential as a flexible solution for broad graph scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。