构建可控制语言类型的词库生成框架,兼顾发音合理性与语义结构。
A Modular Architecture for Typologically Controlled Lexicon Generation

- 模块化设计:从PHOIBLE采音素,用不同语法生成词形,匹配语义本体。
- 概率语法在100-5000词规模下,比确定性和随机基线更优。
- 适合语言学研究、人工语言设计者及需要可控词库的NLP应用。
在计算语言学中,构建发音可读、类型学上合理且语义结构清晰的人工词库仍是未解难题。现有构词工具或缺乏形式化的音系约束,或依赖不透明、不可复现的大模型流水线。本文提出一种模块化框架:从PHOIBLE采样音素库,通过可替换的音系语法(确定性、最优理论、最大熵)生成词形,并基于Swadesh-Leipzig-Jakarta本体实现显式的形式-意义对齐。在100至5000词规模下,通过字符n-gram困惑度、对数似然和与PHOIBLE的KL散度评估表明,概率语法在音系一致性和类型学真实性上均持续优于确定性和随机基线。
原文摘要 · Abstract (English)
Constructing artificial lexicons that are pronounceable, typologically plausible, and semantically structured remains an open challenge in computational linguistics. Existing conlang generators either lack formal phonotactic guarantees or delegate generation to opaque, non-reproducible LLM-based pipelines. We propose a modular framework that samples phoneme inventories from PHOIBLE, generates word forms under interchangeable phonological grammars (deterministic, OT, and MaxEnt), and assigns meanings via a Swadesh--Leipzig--Jakarta ontology with explicit form--meaning alignment. Evaluation on character $n$-gram perplexity, log-likelihood, and KL divergence against PHOIBLE across lexicon sizes of 100-5,000 forms shows that probabilistic grammars consistently outperform deterministic and random baselines on both phonotactic coherence and typological realism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。