arXiv:2508.16243cs.CL2025-08

让小模型学会金融土耳其语,兼顾隐私与专业性

TULIP: Adapting Open-Source Large Language Models for Underrepresented Languages and Specialized Financial Tasks

  • 分五阶段微调小模型,融合真实与合成数据
  • 在土耳其金融任务上准确率显著提升,适配低资源语言
  • 适合关注隐私、垂直领域的小机构或研究者使用

随着大语言模型的普及,其在金融领域的应用潜力巨大。尽管大型专有模型表现优异,但它们作为黑箱API提供,难以满足敏感信息管理与领域知识应用的需求。相比之下,可本地部署的小模型更具灵活性与隐私保障。尤其在低资源语言场景下,如金融领域中的土耳其语,增强小模型能力尤为关键。本文提出TULIP模型,针对Llama 3.1 8B与Qwen 2.5 7B进行领域与语言适应,聚焦土耳其金融应用场景。采用五阶段开发流程:数据收集、持续预训练(CPT)、基准设计、合成数据生成与监督微调(SFT)。实验表明,模型在特定领域和语言任务上的性能得到显著提升。

原文摘要 · Abstract (English)

Thanks to the growing popularity of large language models over the years, there is great potential for their applications in finance. Despite the exceptional performance of larger proprietary models, which are presented as black-box solutions through APIs, smaller models that can be hosted on-premise present opportunities for adaptability and privacy. Especially in cases where the management of sensitive information and application of domain knowledge is important, like finance, enhancing the capabilities of smaller models becomes crucial, notably for underrepresented languages. In this work, we introduce TULIP models, which adapt Llama 3.1 8B and Qwen 2.5 7B for domain and language adaptation, focusing on financial Turkish use cases. The five-stage development pipeline involves data collection, continual pre-training (CPT), benchmark design, synthetic data generation and supervised fine-tuning (SFT). The results show that the capabilities of the models can be enhanced to effectively accomplish targeted tasks in this specific domain and language.

金融AI低资源语言模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。