arXiv:2606.26112cs.CLcs.AI2026-06

用词典构建低资源语言对话系统,效果优于通用模型。

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

论文配图:From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages
图 1 · 摘自论文原文
  • 将词典转化为指令-回复对,构建专用对话数据
  • 120亿参数模型经4比特量化微调,性能达91.0分
  • 适合有词典但无语料的低资源语言开发者

低资源语言在人工智能发展上面临核心挑战:缺乏大规模训练语料库难以构建专业对话系统。本文提出一种系统化方法,将结构化语言资源转化为专用AI系统,证明专家标注的词汇数据库可作为对话AI的有效基础。我们通过将印地语WordNet转换为125万条多样化的指令-回复对,利用4比特量化与高效LoRA技术微调120亿参数语言模型。通过印地语语言学习聊天机器人评估发现,基于结构化知识的系统在教学有效性上表现更优(91.0分),显著高于通用模型的79.4至83.6分,同时保持良好语义表现和极高一致性。该完整流程为具有WordNet资源的任何语言提供了专用AI开发的可行性范例。本研究填补了低资源语言AI可及性的关键空白,提供了一种替代语料密集型方法的实用路径,有望推动数百种具备现有WordNet资源的语言实现专业化AI开发。

原文摘要 · Abstract (English)

Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as effective foundations for conversational AI development. Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization. Evaluation through a Hindi language learning chatbot demonstrates that structured-knowledge-based systems achieve superior pedagogical effectiveness (91.0 vs. 79.4-83.6 for general-purpose models) while maintaining competitive semantic performance and exceptional consistency. The complete pipeline demonstrates a proof-of-concept methodology using Hindi for developing specialized AI systems for any languages with WordNet resources. This work addresses the critical gap in AI accessibility for low-resource languages, offering a practical alternative to corpus-intensive approaches and potentially enabling specialized AI development for the hundreds of languages with existing WordNet resources.

低资源语言对话系统词典构建知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。