提出双路径架构,融合声学与语言信息提升说话风格识别准确率
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
- 设计串行-并行双路径结构,分别处理时序依赖与跨模态交互
- 在8种风格上准确率提升30.3%,参数量减少88.4%
- 适合语音风格分析、多模态语音识别相关研究者使用
说话风格识别(SSR)旨在从语音中识别说话者的风格特征。现有方法主要依赖语言信息,对声学信息整合不足,限制了识别精度的提升。声学与语言模态的融合具有显著提升性能的潜力。本文提出一种新型的串行-并行双路径架构用于SSR,充分利用声学-语言双模态信息。串行路径遵循ASR+STYLE的串行范式,体现时序依赖关系;并行路径引入设计的声学-语言相似性模块(ALSM),实现跨模态的时序同步交互。相比现有基线模型OSUM,本方法参数量减少88.4%,在测试集上对8种说话风格的识别准确率提升30.3%。
原文摘要 · Abstract (English)
Speaking Style Recognition (SSR) identifies a speaker's speaking style characteristics from speech. Existing style recognition approaches primarily rely on linguistic information, with limited integration of acoustic information, which restricts recognition accuracy improvements. The fusion of acoustic and linguistic modalities offers significant potential to enhance recognition performance. In this paper, we propose a novel serial-parallel dual-path architecture for SSR that leverages acoustic-linguistic bimodal information. The serial path follows the ASR+STYLE serial paradigm, reflecting a sequential temporal dependency, while the parallel path integrates our designed Acoustic-Linguistic Similarity Module (ALSM) to facilitate cross-modal interaction with temporal simultaneity. Compared to the existing SSR baseline -- the OSUM model, our approach reduces parameter size by 88.4% and achieves a 30.3% improvement in SSR accuracy for eight styles on the test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。