阿拉伯文字形与功能关系多为任意,随机重映射仍可保持良好模型表现。
Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective
- 用2000种随机字符重映射测试阿拉伯文无点书写下的NLP性能
- 随机重映射在各项任务中表现接近原字符,且降低词汇量和训练成本
- 适合对语言模型轻量化、字符系统设计感兴趣的读者
阿拉伯字母共28个,许多共享基础字形(rasm),仅通过点位区分。早期手稿无点仍可识别,使去点成为检验视觉差异是否必要的自然实验。已有研究显示无点阿拉伯文仍可读且适用于自然语言处理,但其成功是否依赖原始字形分组,或任意但一致的重映射至相同19个无点rasm也能达到相似效果尚不明确。本研究对比标准有点、无点阿拉伯文及约束于同一19个无点rasm的2000种随机字符重映射,采用词级与字符级分词策略,选取四组熵值最高与最低的代表性映射。在语言建模、文本分类、序列标注、机器翻译及原书体还原任务中评估表现。结果表明:保留原始字符区分或传统rasm分组并非强性能所必需。随机重映射在多项任务中表现相当,同时降低词汇量、未登录词率、模型规模与训练开销。研究提示,从NLP视角看,阿拉伯字符的形式-功能关系本质上是任意的,模型更依赖稳定的分布结构而非字形的视觉象形性。
原文摘要 · Abstract (English)
Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dots yet remained interpretable, dot removal offers a natural test of whether these visual distinctions are functionally necessary. Prior work has shown that dotless Arabic can remain readable and effective for natural language processing (NLP), but it remains unclear whether this success depends on preserving the original rasm groupings or whether arbitrary but consistent remappings to the same reduced rasm set can achieve comparable performance. We address this question by comparing standard dotted and dotless Arabic with arbitrary character remappings constrained to the same 19 undotted rasms. We generated 2,000 random remappings under word- and character-level tokenization and selected four representative mappings with the highest and lowest entropy values. These representations were evaluated across language modeling, text classification, sequence labeling, machine translation, and restoration to the original script. The results show that neither preserving original character distinctions nor retaining traditional rasm-based groupings is necessary for strong NLP performance. Random remappings achieve competitive performance while reducing vocabulary size, out-of-vocabulary (OOV) rates, model size, and training cost. These findings suggest that, from an NLP perspective, Arabic character form-function relationships are largely arbitrary: models rely more on stable distributional structure than on the visual iconicity of letter forms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。