arXiv:2510.24992cs.CLeess.AS2025-10ACL被引 12

首个统一处理语音音素任务的模型,支持音频、文本与音素间无缝转换。

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

论文配图:POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
图 1 · 摘自论文原文
  • 构建统一框架,联合处理音素相关任务
  • 在同等规模下超越或持平专用音素模型性能
  • 适合低资源语音处理与通用语音系统研究者

近期语音处理进展在自动语音识别(ASR)、音素识别(PR)、音素转字符(G2P)和字符转音素(P2G)等音素任务上取得显著进步。尽管这些任务概念相似,但长期被孤立研究,依赖各自专用架构与数据集。本文提出POWSM(Phonetic Open Whisper-style Speech Model),首个可联合执行多项音素相关任务的统一框架。该模型实现音频、文本(字符)与音素间的无缝转换,为通用与低资源语音处理开辟新可能。在与同规模专用模型Wav2Vec2Phoneme和ZIPA对比中,性能相当或更优,同时支持G2P、P2G与ASR。训练数据、代码与模型已公开,推动开放科学。

原文摘要 · Abstract (English)

Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G). Despite their conceptual similarity, these tasks have largely been studied in isolation, each relying on task-specific architectures and datasets. In this paper, we introduce POWSM (Phonetic Open Whisper-style Speech Model), the first unified framework capable of jointly performing multiple phone-related tasks. POWSM enables seamless conversion between audio, text (graphemes), and phones, opening up new possibilities for universal and low-resource speech processing. Our model outperforms or matches specialized PR models of similar size (Wav2Vec2Phoneme and ZIPA) while jointly supporting G2P, P2G, and ASR. Our training data, code and models are released to foster open science.

语音处理音素建模多任务学习低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。