arXiv:2509.21698cs.CL2025-09

构建首个金融风险披露无监督主题发现基准,提升投资监管可信度。

GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures

  • 用模型注意力+关键词+共现匹配自动标注161万句金融风险文本
  • 覆盖8247份财报,21个细粒度风险类别,支持标准化评估
  • 适合研究金融文本分析、风险识别与模型可复现性的学者

财务10-K文件中的风险分类对监管和投资决策至关重要,但目前缺乏公开的基准来评估无监督主题模型在该任务上的表现。我们提出GRAB,一个面向金融领域的基准数据集,包含来自8,247份财报的161万条句子,并通过结合FinBERT词元注意力、YAKE关键短语信号与分类法感知的共现匹配,实现无需人工标注的句级标签生成。标签锚定于一个风险分类体系,将193个术语映射到21个细粒度类别,这些类别嵌套在5个宏观类别之下;21个细粒度类型用于弱监督,评估则在宏观层级报告。GRAB统一了评估标准,采用固定数据划分与鲁棒指标——准确率、宏平均F1、主题BERTScore及基于熵的“有效主题数”。该数据集、标签与代码支持经典、基于嵌入、神经及混合主题模型在金融披露文本上的可复现、标准化比较。

原文摘要 · Abstract (English)

Risk categorization in 10-K risk disclosures matters for oversight and investment, yet no public benchmark evaluates unsupervised topic models for this task. We present GRAB, a finance-specific benchmark with 1.61M sentences from 8,247 filings and span-grounded sentence labels produced without manual annotation by combining FinBERT token attention, YAKE keyphrase signals, and taxonomy-aware collocation matching. Labels are anchored in a risk taxonomy mapping 193 terms to 21 fine-grained types nested under five macro classes; the 21 types guide weak supervision, while evaluation is reported at the macro level. GRAB unifies evaluation with fixed dataset splits and robust metrics--Accuracy, Macro-F1, Topic BERTScore, and the entropy-based Effective Number of Topics. The dataset, labels, and code enable reproducible, standardized comparison across classical, embedding-based, neural, and hybrid topic models on financial disclosures.

金融文本主题发现无监督学习风险分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。