135M参数模型高效适配孟加拉语,性能媲美百亿级模型
Surpassing Scale by Efficiency: A Compact 135M Parameter Foundational LLM Natively Adapted for the Bangla Language
- 采用确定性分词合并策略,解决孟加拉语子词碎片化问题
- 零样本多任务测试中超越270M模型,达到10亿参数级表现
- 适合边缘设备、移动端的本地部署,为低资源语言赋能
尽管自然语言处理领域以数十亿参数模型为主导,但在低资源非拉丁文字系统中的部署仍因计算成本过高而受限于边缘设备、移动系统及去中心化本地硬件。本文提出bangla-smollm-135m,一个专为孟加拉语设计的13500万参数解码器型基础模型,通过在TituLLMs与SmolLM2-135M之间采用确定性交集拼接分词合并策略,在不破坏预训练参数状态的前提下,有效缓解了子词文本碎片化问题。在零样本多任务基准测试(PIQA_bn、OpenBookQA_bn、CommonsenseQA_bn和Bangla_MMLU)中,该模型性能匹配甚至超越两倍于其规模的Gemma-3-270m,并达到10亿参数级别模型的水平。模型已开源:rnnandi/bangla-smollm-135m。
原文摘要 · Abstract (English)
While the NLP landscape is dominated by multi-billion parameter architectures, their deployment in low-resource, non-Latin scripts remains computationally prohibitive for edge configurations, mobile systems, and decentralized local hardware. This paper presents bangla-smollm-135m, a highly compact 135-million parameter decoder-only foundational model engineered explicitly for high-efficiency language modeling in the Bangla script. By leveraging a deterministic intersect-and-append token merging strategy between TituLLMs and SmolLM2-135M, the model overcomes subword script fragmentation without destabilizing early pretrained parameter states. In zero-shot multi-task benchmark evaluations (PIQA_bn, OpenBookQA_bn, CommonsenseQA_bn, and Bangla_MMLU), bangla-smollm-135m matches or outperforms models twice its size (Gemma-3-270m) and achieves parity with models in the 1B parameter tier. The model is available at rnnandi/bangla-smollm-135m
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。