arXiv:2603.21571cs.CL2026-03中稿 · presentation at LR…

构建首个英-塔什利赫特平行语料库,助力柏柏尔语正字法标准化与低资源语言处理。

DATASHI: A Parallel English-Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language Processing

  • 设计双版本语料:专家标准版与用户生成非标准版,支持正字法多样性研究。
  • 基于5000句对评估大模型,Gemini-2.5-Pro在词与字符级错误率上最优。
  • 细粒度分析编辑操作,揭示模型对塔什利赫特语音特征的敏感性差异。

DATASHI 是一个全新的英-塔什利赫特平行语料库,填补了阿马齐格语言计算资源的关键空白。该语料库包含5000个句子对,其中1500句具备专家标准化版本与非标准用户生成版本,支持对正字法多样性和规范化进行系统研究。这一双重设计可支撑分词、翻译、正字法归一化等文本类NLP任务,并为读-声数据采集与多模态对齐提供基础。使用前沿大语言模型(GPT-5、Claude-Sonnet-4.5、Gemini-2.5-Pro、Mistral、Qwen3-Max)进行全面评估,显示从零样本到少样本提示均有显著提升,其中Gemini-2.5-Pro在词和字符级错误率上最低,且表现出强跨语言泛化能力。对音位类别(成音节、重音音、软腭音、咽音)中删除、替换、插入等编辑操作的细粒度分析,进一步揭示模型对塔什利赫特显著语音特征的敏感性差异,为低资源柏柏尔语正字法归一化提供新诊断视角。

原文摘要 · Abstract (English)

DATASHI is a new parallel English-Tashlhiyt corpus that fills a critical gap in computational resources for Amazigh languages. It contains 5,000 sentence pairs, including a 1,500-sentence subset with expert-standardized and non-standard user-generated versions, enabling systematic study of orthographic diversity and normalization. This dual design supports text-based NLP tasks - such as tokenization, translation, and normalization - and also serves as a foundation for read-speech data collection and multimodal alignment. Comprehensive evaluations with state-of-the-art Large Language Models (GPT-5, Claude-Sonnet-4.5, Gemini-2.5-Pro, Mistral, Qwen3-Max) show clear improvements from zero-shot to few-shot prompting, with Gemini-2.5-Pro achieving the lowest word and character-level error rates and exhibiting robust cross-lingual generalization. A fine-grained analysis of edit operations - deletions, substitutions, and insertions - across phonological classes (geminates, emphatics, uvulars, and pharyngeals) further highlights model-specific sensitivities to marked Tashlhiyt features and provides new diagnostic insights for low-resource Amazigh orthography normalization.

正字法低资源语言柏柏尔语平行语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。