arXiv:2608.04186cs.CL2026-08

用大模型构建塔吉克语电子详解词典,填补低资源语言工具空白。

Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

论文配图:Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language
图 1 · 摘自论文原文
  • 融合形态分析与大模型生成,构建塔吉克语词典架构。
  • 采用子词分词和参数高效微调,适配标注数据少的现实条件。
  • 首次提出整合传统词典学与生成式AI的完整概念框架。

本文提出一种基于大语言模型(LLMs)的塔吉克语电子详解词典的概念框架。研究动机源于塔吉克语缺乏功能上可比高资源语言的综合性数字词典,且现代自然语言处理技术在低资源语言系统中应用有限。基于对现有语言学、统计及语料资源的系统调研,我们设计了一种集成形态分析、词形还原、语义聚类与词目生成模块的词典架构。子词分词策略的选择由塔吉克语黏着性形态结构及高度变异性决定,并采用适合小规模标注数据的参数高效微调(PEFT)方法。该工作的创新在于首次提出一个统一经典词典学方法、语言统计与大模型生成能力的塔吉克语解释性词典整体架构。其实际意义在于为开发功能完备的电子词典奠定方法论基础,该词典不仅可作为词典工具,还可作为机器翻译、自动摘要、情感分析等任务的核心资源。本文面向计算语言学、词典学及低资源语言NLP系统开发者。

原文摘要 · Abstract (English)

This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lexicographic resource for Tajik that is comparable in functionality to dictionaries for high-resource languages, and from the limited adaptation of modern natural language processing technologies to low-resource language systems. Based on a systematic survey of existing linguistic, statistical, and corpus resources, we propose a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using LLMs. The choice of subword tokenization is justified by the agglutinative nature of Tajik morphology and its high morphological variability, along with a parameter-efficient fine-tuning (PEFT) strategy suitable for limited annotated data. The novelty of the work lies in proposing the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system. The practical significance of the study is the formation of a methodological foundation for developing a full-featured electronic dictionary that can serve both as a lexicographic tool and as a core resource for machine translation, automatic summarization, sentiment analysis, and other applied NLP tasks. The paper is intended for specialists in computational linguistics, lexicography, and developers of natural language processing systems working with low-resource languages.

词典构建大模型低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。