arXiv:2505.13657cs.CLcs.IT2025-05被引 1

用压缩率衡量拼写与发音的对应关系,统一评估不同文字系统的拼写透明度。

Clarifying orthography: Orthographic transparency as compressibility

  • 基于算法信息论,用拼写与发音字符串的互压缩性定义透明度
  • 22种语言验证,结果符合对各文字系统透明度的普遍认知
  • 方法通用且可量化,适合语言学、NLP研究者使用

拼写透明度——拼写与发音之间的直接对应关系——缺乏统一的跨文字体系度量标准。本文借鉴算法信息论思想,将拼写透明度定义为拼写串与发音串之间的互压缩性。该度量能同时捕捉拼写不规则性和规则复杂性两个降低透明度的因素,实现统一量化。通过神经序列模型计算的预序编码长度估算该透明度指标,对22种涵盖字母、辅音音素、元音附标、音节和意音文字等多种书写系统的语言进行评估,结果与人类对各类文字透明度的直观认知一致。互压缩性提供了一种简单、严谨且通用的拼写透明度衡量标准。

原文摘要 · Abstract (English)

Orthographic transparency -- how directly spelling is related to sound -- lacks a unified, script-agnostic metric. Using ideas from algorithmic information theory, we quantify orthographic transparency in terms of the mutual compressibility between orthographic and phonological strings. Our measure provides a principled way to combine two factors that decrease orthographic transparency, capturing both irregular spellings and rule complexity in one quantity. We estimate our transparency measure using prequential code-lengths derived from neural sequence models. Evaluating 22 languages across a broad range of script types (alphabetic, abjad, abugida, syllabic, logographic) confirms common intuitions about relative transparency of scripts. Mutual compressibility offers a simple, principled, and general yardstick for orthographic transparency.

语言学文本生成信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。