arXiv:2507.06795cs.CLcs.AI2025-07EMNLP

用持续预训练让小模型高效适配工业场景,性能提升显著。

ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining

  • 基于领域自适应持续预训练,动态优化小模型
  • 在多个领域实现目标性能大幅提升,通用能力不退化
  • 适合资源有限但需定制化大模型的企业用户

开源大语言模型的兴起为企事业应用带来了新机遇,但许多组织缺乏部署和维护大规模模型的基础设施。因此,小语言模型(sLLMs)成为实用替代方案,尽管存在性能局限。尽管领域自适应持续预训练(DACP)已被探索用于领域适配,其在商业环境中的潜力仍待验证。本研究在多种基础模型和业务领域中验证了基于DACP的方案有效性,构建出ixi-GEN系列sLLMs。通过大量实验与真实场景评估,证明ixi-GEN模型在保持通用能力的同时,显著提升了目标领域的表现,提供了一种成本低、可扩展的企业级部署方案。

原文摘要 · Abstract (English)

The emergence of open-source large language models (LLMs) has expanded opportunities for enterprise applications; however, many organizations still lack the infrastructure to deploy and maintain large-scale models. As a result, small LLMs (sLLMs) have become a practical alternative despite inherent performance limitations. While Domain Adaptive Continual Pretraining (DACP) has been explored for domain adaptation, its utility in commercial settings remains under-examined. In this study, we validate the effectiveness of a DACP-based recipe across diverse foundation models and service domains, producing DACP-applied sLLMs (ixi-GEN). Through extensive experiments and real-world evaluations, we demonstrate that ixi-GEN models achieve substantial gains in target-domain performance while preserving general capabilities, offering a cost-efficient and scalable solution for enterprise-level deployment.

小模型持续预训练工业应用领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。