arXiv:2605.23924cs.CLcs.IR2026-05被引 8

用大模型从财报中提取完整段落信息,提升跨公司比较能力

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

论文配图:Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach
图 1 · 摘自论文原文
  • 基于大模型直接解析10-K文件中的段落披露信息
  • 准确提取报告与嵌套段落数据,支持纵向与横向对比
  • 适合做财务分析、投资者研究及政策评估的学者与从业者

分部披露是财务报告的核心内容,揭示企业内部结构与经济活动在各经营单元间的分配。然而,分部信息常以定性与定量形式分散于10-K文件的表格和文字部分,依赖结构化数据库的实证研究面临完整性与可比性挑战:部分企业年度数据缺失、嵌套分部未被捕捉,且缺乏跨期与跨企业比较支持。本文构建基于大语言模型的框架,直接从10-K文件中提取分部披露信息,并保留可报告与嵌套分部内容。进一步设计检索增强系统,整合多份文件信息以支持可比性分析。通过两个典型场景验证:企业内部纵向分析以解读分部变动趋势,以及不同报告结构企业间地理分部的横向对齐。结果表明,该方法能准确提取分部信息,有效解决需跨期知识的问题,展现了大模型在提升分部披露测量与解读方面的潜力。

原文摘要 · Abstract (English)

Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of economic activities across operating units. However, segment information is often presented in both qualitative and quantitative forms, dispersed across tables and narrative sections of Form 10-K filings. Empirical research relying on structured databases faces both completeness and comparability challenges, as some firm-year observations may be missing, nested segment disclosures are not captured, and support for longitudinal and cross-firm comparability is limited. This study develops a large language model-based framework to extract segment disclosures directly from Form 10-K filings and to preserve both reportable and nested segment information. We further design a retrieval augmented system that incorporates information across multiple filings to support comparability. We use two representative settings to demonstrate its application: longitudinal analysis within a firm to interpret segment changes over time, and cross firm alignment of geographic segments across firms with different reporting structures. The results indicate that the artifact accurately extracts segment-level information and effectively addresses questions that require cross-period knowledge, demonstrating the potential of LLM-based approaches to enhance the measurement and interpretation of segment disclosures.

财务分析大模型信息披露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。