给图表加坐标网格,能显著提升大模型提取数据的准确性。
Spatial Priming Outperforms Semantic Prompting: A Grid-Based Approach to Improving LLM Accuracy on Chart Data Extraction
- 在图表上叠加坐标网格,提供显式空间信息。
- 错误率从25.5%降至19.5%,效果显著(p<0.05)。
- 适合需要高精度图表数据提取的研究者使用。
自动化提取科学图表中的数据是大规模文献分析的关键任务。尽管多模态大语言模型(LLMs)展现出潜力,但在非标准化图表上的准确率仍不理想。本文探讨了两种策略:高层语义提示与低层空间提示。实验发现,基于元数据的两阶段框架和思维链方法未能带来统计显著提升。相比之下,提出一种简单有效的方法:在分析前向图表图像叠加坐标网格。在合成数据集上的定量实验表明,该方法使数据提取误差(SMAPE)从25.5%降至19.5%(p < 0.05),显著优于基线。结论是,对于当前多模态模型,提供明确的空间上下文比高层语义引导更有效可靠。
原文摘要 · Abstract (English)
The automated extraction of data from scientific charts is a critical task for large-scale literature analysis. While multimodal Large Language Models (LLMs) show promise, their accuracy on non-standardized charts remains a challenge. This raises a key research question: what is the most effective strategy to improve model performance (high-level semantic priming) or low-level spatial priming? This paper presents a comparative investigation into these two distinct strategies. We describe our exploratory experiments with semantic methods, such as a two-stage metadata-first framework and Chain-of-Thought, which failed to produce a statistically significant improvement. In contrast, we present a simple but highly effective spatial priming method: overlaying a coordinate grid onto the chart image before analysis. Our quantitative experiment on a synthetic dataset demonstrates that this grid-based approach provides a statistically significant reduction in data extraction error (SMAPE reduced from 25.5% to 19.5%, p < 0.05) compared to a baseline. We conclude that for the current generation of multimodal models, providing explicit spatial context is a more effective and reliable strategy than high-level semantic guidance for this class of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。