arXiv:2502.09661cs.SDeess.AS2025-02

自动提取印度语言的语调特征,助力语音研究与应用。

AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages

  • 基于音节级标注设计工具,适配印度语言的节奏特性。
  • 在泰米尔语、印地语、印英语上准确率接近人工标注。
  • 适合语音合成、语音识别及跨语言研究者使用。

语音信号中的韵律信息在众多应用中具有重要意义,但手动提取耗时费力。为此,本文提出一种名为SIToBI(Segmentation with Intensity, Tones and Break Indices)的工具,可为给定语音信号生成对应的韵律标注,尤其针对印度语言。该工具提供时间对齐的音素、音节和词级转录,音节级基频轮廓、停顿指数及音节级相对强度指数。由于印度语言多为音节计时型,工具重点聚焦音节级标注。尽管当前仅针对泰米尔语、印地语和印英语,但可轻松扩展至其他印度语言乃至其他音节计时语言。通过与人工标注对比评估,工具表现良好。

原文摘要 · Abstract (English)

The availability of prosodic information from speech signals is useful in a wide range of applications. However, deriving this information from speech signals can be a laborious task involving manual intervention. Therefore, the current work focuses on developing a tool that can provide prosodic annotations corresponding to a given speech signal, particularly for Indian languages. The proposed Segmentation with Intensity, Tones and Break Indices (SIToBI) tool provides time-aligned phoneme, syllable, and word transcriptions, syllable-level pitch contour annotations, break indices, and syllable-level relative intensity indices. The tool focuses more on syllable-level annotations since Indian languages are syllable-timed. Indians, regardless of the language they speak, may exhibit influences from other languages. As a result, other languages spoken in India may also exhibit syllable-timed characteristics. The accuracy of the annotations derived from the tool is analyzed by comparing them against manual annotations and the tool is observed to perform well. While the current work focuses on three languages, namely, Tamil, Hindi, and Indian English, the tool can easily be extended to other Indian languages and possibly other syllable-timed languages as well.

韵律分析语音标注印度语言音节级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。