评测文字生成科学图的结构准确性,发现现有模型常有形似神不似问题。
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
- 通过逆向解析生成图像还原结构图,以结构可恢复性为核心评估标准。
- 在复杂拓扑图上,现有模型结构错误率超60%,显示结构保持仍是难题。
- 适合关注科学可视化、结构感知生成的科研与工程人员参考。
科学图表蕴含明确的结构信息,但现代文本到图像模型常生成视觉逼真却结构错误的结果。现有评测或依赖图像中心或主观指标,对结构不敏感;或仅评估中间符号表示,而非最终渲染图像,导致像素级图表生成缺乏有效评估。我们提出SciFlow-Bench,一个以结构为先的基准,直接从像素级输出评估科学图表生成。该数据集源自真实科学PDF,每张源框架图均配以标准真值图,并通过闭环往返协议评估模型:将生成的图表图像逆向解析回结构图进行对比。此设计强调结构可恢复性而非仅视觉相似性,由分层多智能体系统协同规划、感知与结构推理实现。实验表明,保持结构正确性仍是根本挑战,尤其在复杂拓扑场景中,凸显结构感知评估的必要性。
原文摘要 · Abstract (English)
Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results. Existing benchmarks either rely on image-centric or subjective metrics insensitive to structure, or evaluate intermediate symbolic representations rather than final rendered images, leaving pixel-based diagram generation underexplored. We introduce SciFlow-Bench, a structure-first benchmark for evaluating scientific diagram generation directly from pixel-level outputs. Built from real scientific PDFs, SciFlow-Bench pairs each source framework figure with a canonical ground-truth graph and evaluates models as black-box image generators under a closed-loop, round-trip protocol that inverse-parses generated diagram images back into structured graphs for comparison. This design enforces evaluation by structural recoverability rather than visual similarity alone, and is enabled by a hierarchical multi-agent system that coordinates planning, perception, and structural reasoning. Experiments show that preserving structural correctness remains a fundamental challenge, particularly for diagrams with complex topology, underscoring the need for structure-aware evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。