arXiv:2605.20043cs.CL2026-05被引 1

分析日语过去时形态生成中的书写错误,发现80%错误与音节重复有关。

Mind Your Moras: Orthography-Aware Error Analysis of Neural Japanese Morphological Generation

  • 将平假名视为表意系统,分析其对模型泛化的影响
  • 75-80%错误集中在需促音的词干以e结尾的动词上
  • 适用于研究日语形态学或神经模型误差的学者

我们对日语过去时形态屈折进行了一项基于书写的错误分析,将平假名不仅视为转写媒介,更作为编码形态音位差异的表征系统,可能影响模型泛化。在遵循SIGMORPHON 2020和2023共享任务规范的数据集上,评估了两种字符级序列到序列架构在过去时生成上的表现。尽管整体准确率较高,模型仍表现出系统性、语言可解释的错误,这些错误集中在特定平假名书写属性上。我们提出一个包含七种主要失败模式的简洁错误分类体系,并进行了定量与定性分析。促音相关错误占据残差错误的75-80%,尤其在词干以元音「e」结尾且需在过去时后缀前添加促音的动词中更为明显。错误模式在不同架构和随机种子下高度一致,表明书写表征、形态结构与数据频次效应之间存在稳健交互作用。结果强调了在形态复杂语言中开展书写感知评估的重要性。

原文摘要 · Abstract (English)

We present an orthography-aware error analysis of Japanese past-tense morphological inflection, treating hiragana not merely as a transcriptional medium, but as a representational system encoding morphophonological distinctions that may influence model generalization. We evaluate two character-level sequence-to-sequence architectures on past-tense formation using datasets formatted according to the SIGMORPHON 2020 and 2023 shared task conventions. Despite high aggregate accuracy, models exhibit systematic, linguistically interpretable errors that cluster around specific orthographic properties of hiragana. We introduce a concise error taxonomy capturing seven primary failure modes and provide both quantitative and qualitative analyses. Gemination-related errors dominate residual failures, accounting for 75-80% of errors, particularly in verbs whose stems end in the vowel e and require gemination before the past-tense suffix. Error patterns remain highly consistent across architectures and random seeds, suggesting a robust interaction between orthographic representation, morphological structure, and data frequency effects in shaping model generalization. These results underscore the necessity of orthography-aware evaluation for understanding neural generalization in morphologically complex languages.

日语形态学错误分析平假名神经模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。