用大模型精准提取图表拓扑结构,解决视觉与推理双重难题。
TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models

- 分步融合感知与全局结构先验,逐步构建图结构
- 在180张图表上实现边缘识别准确率显著提升
- 适合需要结构化理解的科研与工程场景
图示到图拓扑提取旨在从结构化图示中提取实体及其连接关系构成的图。该任务对现有视觉语言模型仍具挑战性,因其需兼具细粒度感知定位与全局一致性拓扑推理。本文提出TopoBench-180,一个由人工验证的图示到图拓扑提取基准,包含180张涵盖网页型与网络型的结构图,并配有标准图标注。同时提出TopoAgent,一种基于结构感知的感知到推理框架,利用大视觉语言模型实现可靠拓扑提取。TopoAgent通过结合感知定位、全局结构先验、标准节点库构建、以节点为中心的局部到全局关系推理以及拓扑一致性强化,逐步完成目标图提取。在TopoBench-180上的实验表明,TopoAgent优于强基线视觉语言模型及近期视觉推理框架,尤其在边提取方面表现突出。本工作填补了多模态结构理解中的重要空白,建立了图示到图拓扑提取的基准与框架。相关资源将公开发布于https://huggingface.co/datasets/WayneGuo0011/TopoBench-180。
原文摘要 · Abstract (English)
Diagram-to-graph topology extraction aims to extract a graph of entities and their connections from a structural diagram. This task remains challenging for current vision-language models because it requires both fine-grained perceptual grounding and topology-aware reasoning with global consistency. We present TopoBench-180, a human-verified benchmark for diagram-to-graph topology extraction, and TopoAgent, a structure-aware perception-to-reasoning framework for reliable topology extraction using large vision-language models. TopoBench-180 contains 180 structural diagrams spanning Web-style and Network-style categories, paired with canonical graph annotations. TopoAgent progressively extracts the target graph by combining grounded perception, global structural priors, canonical node inventory construction, node-centric local-to-global relation reasoning, and topological consistency enforcement. Experiments on TopoBench-180 show that TopoAgent outperforms strong vision-language model baselines and recent visual reasoning frameworks, especially on edge extraction. More broadly, this work fills an important gap in multimodal structured understanding by establishing a benchmark and framework for diagram-to-graph topology extraction. The benchmark and associated resources will be publicly released at https://huggingface.co/datasets/WayneGuo0011/TopoBench-180.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。