通过动态规划方法生成更粗粒度声学单元,提升长句重合成与低比特率语言建模效果。
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
- 采用惩罚时长的动态规划方法生成粗粒度声学单元。
- 在句子重合成任务中,粗粒度单元显著提升性能;低比特率下词汇与句法建模准确率更高。
- 发现单位粗细并非越粗越好,仅在特定任务中受益,适合语音生成与高效编码场景。
说话人语言模型(SLMs)基于自监督语音表征的离散化声学单元运行。尽管这些单元特性直接影响性能,但码本大小与单元粗粒度(即持续时间)之间的交互关系仍未被探索。本文使用简单的时长惩罚动态规划(DPDP)方法,在不同语言层级上分析了单元粗细对SLM性能的影响。在音素和词级别,只要码本大小合适,粗粒度带来的增益很小。但在整句重合成任务中,粗粒度单元表现更优;在词汇与句法语言建模任务中,粗粒度单元在低比特率下也实现了更高的准确率。因此,我们证明粗粒度单元并非总是更好,但DPDP是一种简单高效的获取有益粗粒度单元的方法。
原文摘要 · Abstract (English)
Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect performance, the interaction between codebook size and unit coarseness (i.e., duration) remains unexplored. We investigate SLM performance as we vary codebook size and unit coarseness using the simple duration-penalized dynamic programming (DPDP) method. New analyses are performed across different linguistic levels. At the phone and word levels, coarseness provides little benefit, as long as the codebook size is chosen appropriately. However, when producing whole sentences in a resynthesis task, SLMs perform better with coarser units. In lexical and syntactic language modeling tasks, coarser units also give higher accuracies at lower bitrates. We therefore show that coarser units aren't always better, but that DPDP is a simple and efficient way to obtain coarser units for the tasks where they are beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。