arXiv:2412.17606cs.CV2024-12被引 4

用分阶段生成法自动生成带密集问答的图表数据集,无需人工标注。

SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images

  • 分阶段生成图表与标注,避免代码错误和重复内容。
  • 生成的图表含完整数据注释和密集问答对,支持高效预训练。
  • 适合需要少量真实数据即实现高性能图表问答模型的研究者。

构建大规模图表问答数据集需大量人力,涉及图表示例收集、属性提取(如文字、数值、颜色)及问答对生成。尽管大语言模型在合成图表方面取得进展,但多数工作仅聚焦于问答生成。直接使用大模型生成图表常出现代码错误、外观相似或内容重复等问题。为此,我们提出SBSFigures(分阶段合成图表)数据集,用于图表问答的预训练。所提流水线可无须人工标注,生成带有完整数据注释和密集问答对的图表。分阶段设计有效提升主题与外观多样性,同时减少代码错误。实验表明,SBSFigures具备强预训练效果,仅需少量真实图表数据即可实现高效微调。

原文摘要 · Abstract (English)

Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors, and generating QAs. Although recent developments in LLMs have led to efforts to synthesize figures, most of these focus primarily on QA generation. Additionally, creating figures directly using LLMs often encounters issues such as code errors, similar-looking figures, and repetitive content in figures. To address this issue, we present SBSFigures (Stage-by-Stage Synthetic Figures), a dataset for pre-training figure QA. Our proposed pipeline enables the creation of chart figures with complete annotations of the visualized data and dense QA annotations without any manual annotation process. Our stage-by-stage pipeline makes it possible to create diverse topic and appearance figures efficiently while minimizing code errors. Our SBSFigures demonstrate a strong pre-training effect, making it possible to achieve efficient training with a limited amount of real-world chart data starting from our pre-trained weights.

图表问答数据合成预训练自动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。