arXiv:2607.15879cs.CLcs.AI2026-07

用大模型自动提取公司治理文件关键信息,准确率高且可标准化。

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

论文配图:DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
图 1 · 摘自论文原文
  • 构建了企业章程与细则的标注数据集,用于评估自动化信息提取。
  • 多数治理条款提取准确率接近上限,但少数条款仍易出错。
  • 优化提示词和流程设计能显著缩小高性能与普通模型差距。

大量实证法律研究依赖将非结构化文本转化为结构化变量。在公司治理研究中,传统做法是人工编码章程与细则,成本高、难扩展且过程不透明。本文提出 DECODEM,一套用于评估从组织文件中自动化提取公司治理变量的基准数据集。该数据集将随机抽取的企业章程和细则与高质量人工标注配对,涵盖实证研究中常见的多种治理条款。利用这些数据集,论文评估了多种基于大语言模型的提取流程,包括不同提示设计、任务拆分与文档处理方式。核心任务为每个治理变量的文档级二分类问题。结果表明,许多条款的自动化提取已达到高精度,中位性能接近各方法的上限。同时,性能在不同变量间系统性差异明显,少数条款贡献了大部分错误。更复杂的提示策略与级联式流程并未一致提升前沿模型表现,但在某些场景下显著缩小了前沿模型与效率导向模型的差距,说明流程设计可在一定程度上弥补模型能力不足。通过提供标准化基准并系统评估提取方法,论文证明当前前沿模型可高精度提取复杂企业文件中的法律意义信息,预示自动化特征提取在未来构建公司治理数据集中的重要作用。

原文摘要 · Abstract (English)

Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from organizational documents. The benchmarks pair randomly sampled corporate charters and bylaws with high-quality human annotations covering a range of governance provisions commonly studied in empirical work. Using these datasets, the paper evaluates several large-language-model extraction pipelines that vary in prompt design, task decomposition, and document handling. The underlying task consists of a set of document-level binary classification problems, one for each governance variable. The results show that automated extraction is feasible at a high level of accuracy for many provisions, with median performance near the upper bound across approaches. At the same time, performance varies systematically across variables, with a small number of provisions accounting for most of the remaining errors. More elaborate prompting strategies and cascading pipelines do not consistently improve performance for frontier models, but substantially narrow the gap between frontier and efficiency-oriented models in some settings, suggesting that pipeline design can partly substitute for model capability. By providing a standardized benchmark and a systematic evaluation of extraction methods, the paper demonstrates that current frontier models can extract legally meaningful information from complex corporate documents with high accuracy and suggests an important future role for automated feature extraction in constructing corporate governance datasets.

信息提取大模型公司治理自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。