用记忆机制提升印尼方言模型训练效率,1.2亿参数三语模型更快更省资源。
Adaptive Engram Memory System for Indonesian Language Model: Generative AI Based on TOBA LM for Batak and Minang Language
- 引入自适应语素记忆系统,通过双三元组路径捕捉词形依赖。
- 训练仅需12,973步,损失从6.4降至1.7996,效率达80%。
- 适合资源有限地区语言建模,尤其支持巴塔克和米南卡保方言。
本研究提出TOBA-LM,一个基于GPT-2架构、含12亿参数的三语语言模型,采用音节聚合分词法,在印尼语、巴塔克语和米南卡保语语料上训练。模型集成恩格拉记忆机制,一种自适应的n-gram记忆系统,拥有50万×768的嵌入表,通过二元组和三元组路径捕捉词形依赖。实证结果显示,训练效率达80%,损失值在仅12,973步内从6.4降至1.7996,远快于传统Transformer架构(需超7万步才达类似收敛)。结果表明,外接统计记忆可显著降低区域语言模型开发的计算需求。
原文摘要 · Abstract (English)
This study presents TOBA-LM, a trilingual language model based on GPT-2 architecture with 1.2 billion parameters, trained on a corpus encompassing Indonesian, Batak, and Minangkabau using syllabic-agglutinative tokenization. The architecture integrates an Engram Memory mechanism, an adaptive n-gram-based memory system with a 500,000 x 768 embedding table that captures morphological dependencies through bigram and trigram pathways. Empirical results demonstrate a training efficiency of 80%, with the loss value dropping from 6.4 to 1.7996 in only 12,973 steps -- significantly faster than the conventional transformer architecture, which required over 70,000 steps to achieve comparable convergence. These findings confirm that the integration of external statistical memory substantially reduces computational requirements for developing regional language models under limited resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。