arXiv:2509.18775cs.CLcs.AI2025-09EMNLP被引 1

用10-K文件自动识别企业间风险关联,提升投资决策效率。

Financial Risk Relation Identification through Dual-view Adaptation

  • 基于10-K文件的时序与词汇模式,无监督微调构建金融领域编码器。
  • 量化生成风险关联分值,准确率显著优于现有方法。
  • 适合金融风控、量化投资研究者使用,代码开源可复现。

众多相互关联的风险事件——从监管变化到地缘政治紧张——可能引发企业间的连锁反应。因此,识别企业间风险关系对投资组合管理与投资策略至关重要。传统方法依赖专家判断和人工分析,主观性强、耗时且难以扩展。为此,我们提出一种系统性方法,利用权威且标准化的10-K文件作为数据源,通过自然语言处理技术捕捉隐含的抽象风险关联。该方法基于文件中的时间序列与词汇模式进行无监督微调,构建具备深层上下文理解能力的领域专用金融编码器,并引入可量化的风险关系评分,实现透明且可解释的分析。大量实验表明,该方法在多个评估场景下均优于强基线模型。代码已公开于 https://github.com/cnclabs/codes.fin.relation。

原文摘要 · Abstract (English)

A multitude of interconnected risk events -- ranging from regulatory changes to geopolitical tensions -- can trigger ripple effects across firms. Identifying inter-firm risk relations is thus crucial for applications like portfolio management and investment strategy. Traditionally, such assessments rely on expert judgment and manual analysis, which are, however, subjective, labor-intensive, and difficult to scale. To address this, we propose a systematic method for extracting inter-firm risk relations using Form 10-K filings -- authoritative, standardized financial documents -- as our data source. Leveraging recent advances in natural language processing, our approach captures implicit and abstract risk connections through unsupervised fine-tuning based on chronological and lexical patterns in the filings. This enables the development of a domain-specific financial encoder with a deeper contextual understanding and introduces a quantitative risk relation score for transparency, interpretable analysis. Extensive experiments demonstrate that our method outperforms strong baselines across multiple evaluation settings. Our codes are available at https://github.com/cnclabs/codes.fin.relation.

金融风险文本挖掘无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。