arXiv:2601.15247cs.CL2026-01被引 1

用AI从财报中精准提取风险信息,并自动优化分类体系。

Taxonomy-Aligned Risk Extraction from 10-K Filings with Autonomous Improvement Using LLMs

  • 三阶段流程:LLM提取+语义映射+AI评分过滤,确保分类准确。
  • 从500强公司提取1.07万条风险,同行业公司风险相似度高63%。
  • AI自动发现分类问题并优化,嵌入分离度提升超100%,适合金融风控场景。

我们提出一种从企业10-K财报中结构化提取风险因素的方法,同时保持与预定义层级分类体系的一致性。该方法采用三阶段流程:先通过大模型提取风险因素及原文佐证,再利用嵌入式语义映射匹配到分类体系,最后由大模型作为裁判对错误分配进行过滤。我们从中提取了10,688条风险因素,分析了不同行业集群间的风险特征相似性。此外,我们引入自主分类体系维护机制:由AI代理分析评估反馈,识别问题类别、诊断失败模式并提出改进方案,在案例研究中实现嵌入分离度104.7%的提升。外部验证表明,同一行业的公司风险特征相似度比跨行业对高出63%(Cohen's d=1.06,AUC 0.82,p<0.001),证明分类体系具有经济意义。该方法可推广至任何需要从非结构化文本中进行分类对齐提取的领域,且通过自主迭代持续提升质量。

原文摘要 · Abstract (English)

We present a methodology for extracting structured risk factors from corporate 10-K filings while maintaining adherence to a predefined hierarchical taxonomy. Our three-stage pipeline combines LLM extraction with supporting quotes, embedding-based semantic mapping to taxonomy categories, and LLM-as-a-judge validation that filters spurious assignments. To evaluate our approach, we extract 10,688 risk factors from S&P 500 companies and examine risk profile similarity across industry clusters. Beyond extraction, we introduce autonomous taxonomy maintenance where an AI agent analyzes evaluation feedback to identify problematic categories, diagnose failure patterns, and propose refinements, achieving 104.7% improvement in embedding separation in a case study. External validation confirms the taxonomy captures economically meaningful structure: same-industry companies exhibit 63% higher risk profile similarity than cross-industry pairs (Cohen's d=1.06, AUC 0.82, p<0.001). The methodology generalizes to any domain requiring taxonomy-aligned extraction from unstructured text, with autonomous improvement enabling continuous quality maintenance and enhancement as systems process more documents.

风险提取LLM应用分类优化财报分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。