arXiv:2508.13404cs.AIcs.CL2025-08Conference of the …

用智能代理自动提取金融表格数据,解决跨页混乱表的结构化难题。

TASER: Table Agents for Schema-guided Extraction and Recommendation

  • 构建多阶段智能体系统,按预设投资框架自动识别、提取和推荐表格内容。
  • 在99.4%无边框的复杂表格上,比传统模型提升10.1%提取准确率。
  • 支持持续学习,批量越大,推荐修正方案越多,有效提取量增9.8%。

真实世界的财务文件包含实体投资持仓的关键信息,对评估风险、盈利和关系网络至关重要。然而这些信息常隐藏于杂乱、跨页、碎片化的表格中,难以解析,阻碍下游问答与数据标准化。我们分析发现,99.4%的财务表格缺乏边界框,最大表格长达44页。为此,我们提出TASER(Table Agents for Schema-guided Extraction and Recommendation),一个持续学习的智能体式表格提取系统,将高度非结构化的多页异构表格转化为符合规范的结构化输出。系统基于初始投资组合模式,统一执行表格检测、分类、提取与推荐。其推荐智能体可审查未匹配结果并提出模式修订建议,使TASER相比视觉表检测模型Table Transformer提升10.1%性能。在持续学习过程中,更大的批量带来104.3%的有用推荐增长和9.8%的总提取量提升。为训练TASER,我们人工标注了22,584页、3,213张表格,覆盖7317亿美元持仓,形成TASERTab数据集,以推动真实金融表格研究。结果表明,持续学习智能体在复杂表格提取中具有显著潜力。

原文摘要 · Abstract (English)

Real-world financial filings report critical information about an entity's investment holdings, essential for assessing that entity's risk, profitability, and relationship profile. Yet, these details are often buried in messy, multi-page, fragmented tables that are difficult to parse, hindering downstream QA and data normalization. Specifically, 99.4% of the tables in our financial table dataset lack bounding boxes, with the largest table spanning 44 pages. To address this, we present TASER (Table Agents for Schema-guided Extraction and Recommendation), a continuously learning, agentic table extraction system that converts highly unstructured, multi-page, heterogeneous tables into normalized, schema-conforming outputs. Guided by an initial portfolio schema, TASER executes table detection, classification, extraction, and recommendations in a single pipeline. Our Recommender Agent reviews unmatched outputs and proposes schema revisions, enabling TASER to outperform vision-based table detection models such as Table Transformer by 10.1%. Within this continuous learning process, larger batch sizes yield a 104.3% increase in useful schema recommendations and a 9.8% increase in total extractions. To train TASER, we manually labeled 22,584 pages and 3,213 tables covering $731.7 billion in holdings, culminating in TASERTab to facilitate research on real-world financial tables and structured outputs. Our results highlight the promise of continuously learning agents for robust extractions from complex tabular data.

表格提取金融数据智能体持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。