2.7B参数德国语模型,专为手机端高效推理设计
From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference

- 针对德语特点优化数据处理与模型架构
- 在55千小时H100算力下达成7B模型性能
- 适合移动端部署的轻量级德语语言模型
我们提出ELMOD——一种专为设备端部署设计的高效语言模型,是一个2.7B参数的德语语言模型,可在资源受限硬件上实现高效推理。该模型在有限计算预算(55,000小时H100 GPU)下,仅使用公开数据训练完成。我们开发了针对德语特点的数据预处理方法,包括对形态变化、复合词和拼写规范的特殊处理,区别于英语主流方案。此外,引入数据质量过滤与重述步骤,提升了数据的教学质量,改善了渐进式训练阶段表现,并降低整体算力需求。得益于模型架构与数据选择的协同优化,包括预过滤与教学品质提升策略,ELMOD在小于30亿参数的模型中表现最强,性能媲美70亿参数模型。
原文摘要 · Abstract (English)
We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (<3B), matching the performance of 7B-parameter models in German.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。