arXiv:2409.14524cs.IRcs.DL2024-09

一键提取PDF表格,支持手动选区,提升数据获取效率。

tabulapdf: An R Package to Extract Tables from PDF Documents

  • 基于Tabula Java库,直接将PDF表格导入R
  • 提供自动与手动两种提取模式,操作灵活
  • 适合需要高效处理表格的科研与调查记者

tabulapdf 是一个 R 包,利用 Tabula Java 库将 PDF 文件中的表格直接导入 R 环境。该工具可显著减少调查性新闻等领域的数据提取耗时与人力成本。支持自动和手动两种表格提取方式,其中手动模式通过 Shiny 界面实现,用户可用鼠标在PDF上直接选择目标区域进行数据抓取。

原文摘要 · Abstract (English)

tabulapdf is an R package that utilizes the Tabula Java library to import tables from PDF files directly into R. This tool can reduce time and effort in data extraction processes in fields like investigative journalism. It allows for automatic and manual table extraction, the latter facilitated through a Shiny interface, enabling manual areas selection with a computer mouse for data retrieval.

PDF处理数据提取R语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。