将Llama 3.1微调至电商领域,构建专用大模型。
Domain Adaptation of Foundation LLMs for e-Commerce
- 在1万亿电商文本上持续预训练Llama 3.1
- 多语言评测显示电商理解能力显著提升
- 支持与基础模型融合,平衡通用与专精性能
我们提出e-Llama系列模型:80亿和700亿参数的大语言模型,专为电商领域设计。这些模型基于Llama 3.1,在1万亿条领域特定数据上进行持续预训练,旨在作为具备深度电商知识的通用基础模型,支持后续指令微调。通过一系列消融实验,我们讨论并验证了超参数选择的有效性。为量化模型在电商领域的适配程度,我们定义并实现了多语言、电商专用评估任务。结果表明,合理设定训练配置后,模型可在不显著牺牲通用任务性能的前提下完成领域适应。此外,我们探索了将适配模型与基底模型融合的可能性,以更灵活地控制跨领域性能权衡。
原文摘要 · Abstract (English)
We present the e-Llama models: 8 billion and 70 billion parameter large language models that are adapted towards the e-commerce domain. These models are meant as foundation models with deep knowledge about e-commerce, that form a base for instruction- and fine-tuning. The e-Llama models are obtained by continuously pretraining the Llama 3.1 base models on 1 trillion tokens of domain-specific data. We discuss our approach and motivate our choice of hyperparameters with a series of ablation studies. To quantify how well the models have been adapted to the e-commerce domain, we define and implement a set of multilingual, e-commerce specific evaluation tasks. We show that, when carefully choosing the training setup, the Llama 3.1 models can be adapted towards the new domain without sacrificing significant performance on general domain tasks. We also explore the possibility of merging the adapted model and the base model for a better control of the performance trade-off between domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。