实现语音任意片段时长修改,保持音高谱等核心特征不变
The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
- 通过指定起止时间或目标时长,灵活调整语音任意片段时长
- 支持多段同时修改,且不改变原始信号的音高轮廓与频谱特性
- 适合语音合成、语音编辑等需要精准控制时长的应用场景
本文深入评估和分析了多种主流算法,这些算法能够在不改变原始语音信号关键属性(如音高轮廓、功率谱等)的前提下,任意修改给定语音信号中任意部分的时长。此处的任意修改指可指定需修改区间的起始与结束时间,或目标时长,目标时长可为时间域的固定值,也可为原时长的缩放因子。此外,任意修改还意味着可在同一时刻修改任意数量的区间。
原文摘要 · Abstract (English)
This paper deeply evaluates and analyzes several mainstream algorithms that can arbitrarily modify the duration of any portion of a given speech signal without changing the essential properties (e.g., pitch contour, power spectrum, etc.) of the original signal. Arbitrary modification in this context means that the duration of any region of the signal can be changed by specifying the starting and ending time for modification or the target duration of the specified interval, which can be either a fixed value of duration in the time domain or a scaling factor of the original duration. In addition, arbitrary modification also indicates any number of intervals can be modified at the same time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。