用大模型生成日文流程图问答数据集,提升视觉语言模型理解能力。
JSynFlow: Japanese Synthesised Flowchart Visual Question Answering Dataset built with Large Language Models
- 通过领域语言代码生成流程图图像,结合大模型合成问答对。
- 在流程图问答任务上,微调后模型准确率显著提升。
- 适合研究日文文档理解、流程图分析与多模态模型训练者使用。
视觉语言模型(VLMs)有望通过问答接口分析包含流程图的复杂文档。识别与解读流程图的能力需求迫切,因其提供了文本无法传达的洞察。然而,构建具备精确流程图理解能力的VLM需要大规模的流程图图像与对应文本数据集,而数据创建过程耗时费力。为解决此问题,我们提出JSynFlow,一个基于大语言模型(LLMs)生成的日文流程图视觉问答数据集。该数据集包含各类职业的任务描述、由领域特定语言(DSL)代码渲染的流程图图像,以及相关的问答对。本文详述了数据集的合成流程,并证明在JSynFlow上微调可显著提升VLM在流程图问答任务上的表现。数据集已公开于https://huggingface.co/datasets/jri-advtechlab/jsynflow。
原文摘要 · Abstract (English)
Vision and language models (VLMs) are expected to analyse complex documents, such as those containing flowcharts, through a question-answering (QA) interface. The ability to recognise and interpret these flowcharts is in high demand, as they provide valuable insights unavailable in text-only explanations. However, developing VLMs with precise flowchart understanding requires large-scale datasets of flowchart images and corresponding text, the creation of which is highly time-consuming. To address this challenge, we introduce JSynFlow, a synthesised visual QA dataset for Japanese flowcharts, generated using large language models (LLMs). Our dataset comprises task descriptions for various business occupations, the corresponding flowchart images rendered from domain-specific language (DSL) code, and related QA pairs. This paper details the dataset's synthesis procedure and demonstrates that fine-tuning with JSynFlow significantly improves VLM performance on flowchart-based QA tasks. Our dataset is publicly available at https://huggingface.co/datasets/jri-advtechlab/jsynflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。