构建首个大规模结构化IPO文档数据集,助力长文本多模态分析
IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

- 开发开源工具链,将超长招股书拆解为标准化章节文本与图像
- 建成覆盖10.9万份文件、7.6万张图表的多模态数据集,时间跨度达32年
- 揭示主流多模态模型在财务图表判读上与人类专家存在显著偏差
首次公开募股(IPO)文件是企业上市时发布的长篇多模态文档,包含文本、图表等信息,对资本市场至关重要。然而,现有研究缺乏大规模、标准化的数据集与评估基准。本文提出IPO-Toolkit,一个开源框架,可自动下载并解析超过50万词的IPO文件,将其按章节结构化,并提取嵌入图像。基于此,构建了IPO-Dataset,涵盖1994至2026年间超过109,000份IPO文件及修订版,含76,000余张图表。该数据集支持对财务图表的质量与误导性进行结构化评估。实验表明,当前先进多模态模型在这些任务上的表现常偏离专家判断,暴露其在处理真实监管文档时的对齐问题。此外,数据集还可用于分析不同行业间披露模式的文本与视觉差异。代码、数据与网站已公开,采用CC-BY-4.0许可。
原文摘要 · Abstract (English)
An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its shares. These filings describe a firm's business, financials, and risks and are long, multimodal documents with narrative text and images. Despite their importance to financial markets, there is no large-scale, standardized dataset or benchmark for studying IPO filings with modern language and multimodal models. These documents pose significant challenges: filings frequently exceed 500,000 tokens and lack consistent structural organization. We introduce the IPO-Toolkit, an open-source framework for downloading and parsing IPO filings into standardized section-structured text and extracted images. The toolkit segments filings, extracts embedded images, and produces structured outputs that enable large-scale, reproducible analysis workflows over long, multimodal documents. Using this infrastructure, we construct the IPO-Dataset, a large, section-structured, multimodal dataset covering more than 109,000 IPO filings and amendments from 1994 to 2026 and containing over 76,000 images. We establish structured evaluation tasks over extracted financial charts, including chart quality and misleadingness assessment. Our experiments show that state-of-the-art multimodal models often diverge from expert human judgments on these tasks, exposing alignment challenges in multimodal reasoning over long, real-world regulatory documents. Beyond benchmarking, the IPO-Dataset enables large-scale analysis of section-level textual variation and cross-industry differences in visual and textual disclosure practices. Our code, dataset, and website are publicly available under CC-BY-4.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。