arXiv:2606.18192cs.AI2026-06

将美国上市公司财报重构为可高效训练的语言模型数据集。

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

论文配图:The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data
图 1 · 摘自论文原文
  • 把证监会文件转成布局忠实的多标记格式,适合金融语言建模。
  • 释放1520亿个标记的公开数据,与主流语料重叠不足0.1%。
  • 适合金融推理、预测与合规分析,支持真实财务文档理解。

随着高质量公开网页语料逐渐枯竭,长文本高质量文档成为大语言模型训练数据的稀缺资源。现有长文本语料多为专有、成本高昂、合成生成或集中在编程等狭窄领域。我们提出斯坦福EDGAR文件数据集(SEFD),将美国证券交易委员会(SEC) filings 重构为布局忠实的 MultiMarkdown 格式,用于金融语言建模与评估。SEFD使经审计的财务报表、风险披露、所有权报告、会计附注及市场重大事件文件可用于长文本预训练,并作为金融推理、预测、合规和文档理解的基础。该语料具备高词元效率、模型就绪性,且与 Common Crawl 衍生语料重叠率低于0.1%。我们发布 SEFD-v1,包含1520亿词元的初始公开快照,并对一个包含1850万份文件的档案进行估算,总词元量达5500亿。此外,我们引入两个基于SEFD的基准:EDGAR-Forecast,评估模型在知识截止后对文件内容的数值预测能力;EDGAR-OCR,评估复杂财务表格的文本识别性能。

原文摘要 · Abstract (English)

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora are often proprietary and costly to acquire, synthetically generated, or concentrated in narrow domains such as programming. We introduce the Stanford EDGAR Filings Dataset (SEFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation. SEFD makes audited financial statements, risk disclosures, ownership reports, accounting notes, and market-moving event filings usable as long-context pretraining data and as a basis for financial reasoning, forecasting, compliance, and document understanding. The resulting corpus is token-efficient, model-ready, and has less than 0.1% overlap with Common Crawl-derived corpora. We release SEFD-v1, a 152B-token initial public snapshot, and provide corpus-level analyses of a larger 18.5M-filing archive estimated at 550B tokens. We further introduce two SEFD-derived benchmarks: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.

金融AI数据重构长文本开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。