NoLBERT用1976-1995文本训练,避免信息泄露,提升经济预测准确性。
NoLBERT: A No Lookahead(back) Foundational Language Model
- 仅用1976-1995年文本预训练,规避前后信息泄露
- 在专利数据上构建企业创新网络,预测长期盈利增长
- 轻量模型适合经济学、金融学等实证研究
我们提出NoLBERT,一种轻量级、带时间戳的基础语言模型,适用于经济学、金融学及社会科学中的实证研究,尤其擅长预测。通过仅在1976至1995年的文本上进行预训练,该模型避免了回溯与前瞻偏差(信息泄露),从而保障计量推断的有效性。在自然语言处理基准测试中,其表现超越领域专用基线,同时保持时间一致性。应用于专利文本时,NoLBERT可构建企业级创新网络,并显示创新中心性提升能有效预测长期利润增长。
原文摘要 · Abstract (English)
We present NoLBERT, a lightweight, timestamped foundational language model for empirical research -- particularly for forecasting in economics, finance, and the social sciences. By pretraining exclusively on text from 1976 to 1995, NoLBERT avoids both lookback and lookahead biases (information leakage) that can undermine econometric inference. It exceeds domain-specific baselines on NLP benchmarks while maintaining temporal consistency. Applied to patent texts, NoLBERT enables the construction of firm-level innovation networks and shows that gains in innovation centrality predict higher long-run profit growth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。