arXiv:2608.15799cs.CL2026-08中稿 · the Proceedings of…

Mimi语音模型的2048个语义词元可对应到不同层级的音素单位。

Using the Mimi codec for metalinguistic representations

论文配图:Using the Mimi codec for metalinguistic representations
图 1 · 摘自论文原文
  • 通过重新对齐语料库,揭示词元与音素层级的映射关系。
  • 实验表明原方法无法捕捉语义词元到发音的映射。
  • 适用于语音合成与音素级表示研究者。

本文聚焦于Moshi语言模型中神经编解码器Mimi所使用的2048个语义词元字典。我们发现,使用Mimi进行的ABX实验无法捕捉语义词元与音位实现之间的映射关系。通过对齐Mimi表示与TIMIT语料库的转录文本,我们证实这2048个词元ID分别对应四音素、三音素、双音素、单音素及子音素等不同层级的发音实现。

原文摘要 · Abstract (English)

In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX experiment carried out with Mimi fails to capture the mapping of the semantic tokens to phone realisations. By realigning Mimi representations to the TIMIT corpus transcriptions, we show that the 2048 tokens IDs of the semantic codebook map to quadphone, triphone, biphone, phone and subphone realisations.

语音编码语义表示音素层级Mimi

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。