构建跨文字塔吉克语-波斯语词表,提升低资源语言处理效率
TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP
- 设计轻量级混合模型,结合规则与检索实现跨文字匹配
- 神经与检索模型达98%-99%准确率,混合模型96.4%用于OCR纠错
- 专为低资源场景优化,适合实际应用与可解释性需求
本文提出TajPersLexon,一个包含40,112个词和短语对的塔吉克语-波斯语并行词汇资源,用于跨文字的词项检索、音译和对齐。我们在仅使用CPU的环境下,对比了三类方法:(i)轻量级混合管道,(ii)神经序列到序列模型,(iii)检索方法。实验表明该任务基本可解,神经与检索基线在顶1准确率上达到98%-99%。关键发现是,大型多语言句子编码器在此词汇匹配任务中表现不佳,而我们提出的可解释混合模型在实际应用中展现出良好的准确率-效率权衡,于OCR后纠错任务中取得96.4%准确率。所有实验均使用固定随机种子以保证可复现性。数据集、代码与模型将公开发布。
原文摘要 · Abstract (English)
This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical retrieval, transliteration, and alignment in low-resource settings. We conduct a comprehensive CPU-only benchmark comparing three methodological families: (i) a lightweight hybrid pipeline, (ii) neural sequence-to-sequence models, and (iii) retrieval methods. Our evaluation establishes that the task is essentially solvable, with neural and retrieval baselines achieving 98-99% top-1 accuracy. Crucially, we demonstrate that while large multilingual sentence transformers fail on this exact lexical matching, our interpretable hybrid model offers a favorable accuracy-efficiency trade-off for practical applications, achieving 96.4% accuracy in an OCR post-correction task. All experiments use fixed random seeds for full reproducibility. The dataset, code, and models will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。