低资源环境下提升大模型领域适应效率,兼顾能效与硬件成本
Low-resource domain adaptation while minimizing energy and hardware resource consumption
- 探索不同数值精度与数据并行策略对训练速度的影响
- 发现混合精度可显著降低能耗且保持模型准确率
- 适合资源受限团队在小规模设备上做领域适配
训练大型语言模型(LLMs)在能源、硬件和标注数据方面成本高昂,常导致模型偏向主流文化与价值观(Santy et al., 2023)。领域适应被视为提升模型与多元文化价值契合度的可行方案(Hershcovich et al., 2022),但其计算开销仍是一大障碍,尤其对缺乏大规模基础设施的研究团队。本文评估了不同数值精度格式与数据并行策略对训练速度(作为能效与硬件消耗的代理指标)及模型准确率的影响,旨在推动低资源环境下的领域适应实践。研究结果适用于能源效率、可及性或硬件资源有限的所有场景。
原文摘要 · Abstract (English)
Training Large Language Models (LLMs) is costly in terms of energy, hardware, and annotated data, often resulting in a positionality rooted in predominant cultures and values (Santy et al., 2023). Domain adaptation has emerged as a promising strategy to better align models with diverse cultural and value contexts (Hershcovich et al., 2022), but its computational cost remains a significant barrier, particularly for research groups lacking access to large-scale infrastructure. In this paper, we evaluate how the use of different numerical precision formats and data parallelization strategies impacts both training speed (as a proxy to energy and hardware consumption) and model accuracy, with the goal of facilitating domain adaptation in low-resource environments. Our findings are relevant to any setting where energy efficiency, accessibility, or limited hardware availability are key concerns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。