通过分层自适应集成微调,用小模型实现金融NLP性能超越大模型。
LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
- 分析隐藏状态选择性微调关键层,冻结非必要层降低计算开销。
- 30亿参数模型在金融任务上超越GPT-4等大模型表现。
- 适合资源有限但需高精度金融文本分析的机构部署使用。
自然语言处理(NLP)正在推动金融行业变革,广泛应用于文本分析、风险管理和预测等领域。像BloombergGPT和FinMA这样的大语言模型(LLM)已在情感分析、股票走势预测和信用风险评估等任务上树立新基准。此外,双语金融大模型FinMA-ES在FLARE和FLARE-ES基准测试中也表现出色。然而,这些模型的高计算需求限制了众多组织的应用。为此,我们提出分层自适应集成微调(LAET),通过分析隐藏状态表示,选择性地微调预训练语言模型中最有效的层,同时冻结次要层。该方法显著降低计算开销,同时提升特定任务性能。实验表明,该方法在金融NLP任务中表现优异,即使使用约30亿参数的小型模型,其效果仍超过现有基准和包括GPT-4在内的先进模型。本研究为前沿金融NLP技术与实际部署之间架起桥梁,提供高效且可扩展的金融应用模型。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) has transformed the financial industry, enabling advancements in areas such as textual analysis, risk management, and forecasting. Large language models (LLMs) like BloombergGPT and FinMA have set new benchmarks across various financial NLP tasks, including sentiment analysis, stock movement prediction, and credit risk assessment. Furthermore, FinMA-ES, a bilingual financial LLM, has also demonstrated strong performance using the FLARE and FLARE-ES benchmarks. However, the high computational demands of these models limit the accessibility of many organizations. To address this, we propose Layer-wise Adaptive Ensemble Tuning (LAET), a novel strategy that selectively fine-tunes the most effective layers of pre-trained LLMs by analyzing hidden state representations while freezing less critical layers. LAET significantly reduces computational overhead while enhancing task-specific performance. Our approach shows strong results in financial NLP tasks, outperforming existing benchmarks and state-of-the-art LLMs such as GPT-4, even with smaller LLMs ($\sim$3B parameters). This work bridges cutting-edge financial NLP research and real-world deployment with efficient and scalable models for financial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。