arXiv:2505.20650cs.CLcs.AI2025-05被引 4

首个面向财务报告结构化标注的综合评测基准,评估大模型在真实场景下的数值与概念匹配能力。

FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information

  • 分两阶段设计:先提取数值实体,再映射到完整美国会计准则分类体系
  • 零样本测试显示模型在概念匹配上表现差,暴露领域结构理解短板
  • 适用于金融信息抽取、大模型财务推理能力评估的研究者

准确解读财务报告中的数值信息对市场和监管至关重要。尽管XBRL提供了财务数据标注标准,但将数千个事实映射到超过10,000个美国会计准则(US GAAP)概念仍成本高且易出错。现有评测将此任务简化为小范围单步分类,忽略分类体系的层级语义和财务文档的结构特性,无法真实评估大语言模型(LLMs)性能。为此,我们提出FinTagging,首个面向结构感知与全范围XBRL标注的综合性评测基准。我们将复杂标注过程分解为两个子任务:(1) FinNI(财务数值识别),从文本和表格等异构上下文中提取实体与类型;(2) FinCL(财务概念链接),将提取的实体映射至完整的美国会计准则分类体系。该两阶段设计使对大模型在数值推理与分类体系对齐能力的公平评估成为可能。在零样本设置下评估多种大模型发现,模型在提取任务上泛化良好,但在细粒度概念链接上表现显著不足,揭示了其在特定领域结构化推理中的关键局限。

原文摘要 · Abstract (English)

Accurate interpretation of numerical data in financial reports is critical for markets and regulators. Although XBRL (eXtensible Business Reporting Language) provides a standard for tagging financial figures, mapping thousands of facts to over 10k US GAAP concepts remains costly and error prone. Existing benchmarks oversimplify this task as flat, single step classification over small subsets of concepts, ignoring the hierarchical semantics of the taxonomy and the structured nature of financial documents. Consequently, these benchmarks fail to evaluate Large Language Models (LLMs) under realistic reporting conditions. To bridge this gap, we introduce FinTagging, the first comprehensive benchmark for structure aware and full scope XBRL tagging. We decompose the complex tagging process into two subtasks: (1) FinNI (Financial Numeric Identification), which extracts entities and types from heterogeneous contexts including text and tables; and (2) FinCL (Financial Concept Linking), which maps extracted entities to the full US GAAP taxonomy. This two stage formulation enables a fair assessment of LLMs' capabilities in numerical reasoning and taxonomy alignment. Evaluating diverse LLMs in zero shot settings reveals that while models generalize well in extraction, they struggle significantly with fine grained concept linking, highlighting critical limitations in domain specific structure aware reasoning.

财务信息抽取大模型评测结构化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。