arXiv:2601.04597cs.CL2026-01

用模型合并技术打造金融+泰语双专精小模型,兼顾性能与成本。

THaLLE-ThaiLLM: Domain-Specialized Small LLMs for Finance and Thai -- Technical Report

  • 通过合并Qwen-8B与ThaiLLM-8B,提升泰语通用能力。
  • 融合金融专用模型后,多领域基准得分均显著提高。
  • 适合需要本地部署、预算有限的金融机构使用。

大型语言模型在银行与金融领域展现出巨大潜力,可实现复杂任务自动化与规模化决策支持。由于隐私、安全与监管顾虑,机构常偏好本地部署。泰国语言模型计划(ThaiLLM)旨在增强开源大模型的泰语能力,推动本土产业应用。然而,部署多个专用模型与训练单一多功能模型之间存在成本权衡。为此,我们探索模型合并作为资源高效的替代方案。实验表明:将Qwen-8B与ThaiLLM-8B合并后,泰语通用能力显著提升,在M3和M6 O-NET考试中得分更高;进一步融合THaLLE-CFA-8B后,通用与金融领域表现同步提升,涵盖M3、M6 O-NET、Flare-CFA及Thai-IC等多基准测试。本报告验证了模型合并在高效构建多能力小模型方面的可行性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant potential across various domains, particularly in banking and finance, where they can automate complex tasks and enhance decision-making at scale. Due to privacy, security, and regulatory concerns, organizations often prefer on-premise deployment of LLMs. The ThaiLLM initiative aims to enhance Thai language capabilities in open-LLMs, enabling Thai industry to leverage advanced language models. However, organizations often face a trade-off between deploying multiple specialized models versus the prohibitive expense of training a single multi-capability model. To address this, we explore model merging as a resource-efficient alternative for developing high-performance, multi-capability LLMs. We present results from two key experiments: first, merging Qwen-8B with ThaiLLM-8B demonstrates how ThaiLLM-8B enhances Thai general capabilities, showing an uplift of M3 and M6 O-NET exams over the general instruction-following Qwen-8B. Second, we merge Qwen-8B with both ThaiLLM-8B and THaLLE-CFA-8B. This combination results in further improvements in performance across both general and financial domains, by demonstrating an uplift in both M3 and M6 O-NET, Flare-CFA, and Thai-IC benchmarks. The report showcases the viability of model merging for efficiently creating multi-capability LLMs.

小模型金融AI泰语模型合并

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。