arXiv:2502.15411cs.CLcs.AI2025-02中稿 · LREC 2026被引 1

构建了百万级财报KPI数据集,支持多任务金融文本分析。

HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings

  • 构建198类分层标签体系,覆盖165万段财报文本
  • 编码器模型分类宏F1超0.906,大模型结构化提取F1达0.440
  • 专攻财报日期识别误差,适合金融信息抽取研究者

准确标注财报内容可为利益相关方带来显著短期收益。公开财务报告强制使用机器可读的iXBRL格式,但其复杂的细粒度分类体系限制了关键绩效指标(KPI)在不同公司间的迁移能力。为此,我们提出层次化金融关键绩效指标(HiFi-KPI)数据集,包含165万段文本和19.8万唯一、分层组织的标签,与iXBRL分类体系对齐。该数据集支持多种任务,我们评估了三项:KPI分类、KPI抽取和结构化KPI抽取。为快速评估,我们还发布了人工精选的8,000段文本子集HiFi-KPI-Lite。在该子集上,基于编码器的模型分类宏F1超过0.906,而大型语言模型(LLMs)在结构化抽取任务中达到0.440的F1。定性分析表明,抽取错误主要源于日期识别问题。所有代码与数据已开源至https://github.com/aaunlp/HiFi-KPI。

原文摘要 · Abstract (English)

Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 8K paragraph subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data at https://github.com/aaunlp/HiFi-KPI.

金融NLPKPI抽取数据集iXBRL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。