首个直接从手语视频生成语音的统一框架,避免文本中间环节误差。
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
- 统一框架直接从手语视频生成语音,不依赖中间文本。
- 引入姿态感知视觉模块和语义对齐池,提升音视频精准映射。
- 适用于听障人士语音重建,尤其适合中文手语语音转换场景。
手语(Cued Speech, CS)通过手势编码增强口读能力,为听力障碍者提供清晰的视觉语音线索。视频到语音生成(CSV2S)旨在将手语视频转化为可理解的语音信号。现有研究多聚焦于手语识别(CSR),常采用先识别再转语音的两阶段流程,但依赖文本中间表示易引发错误传播与时间错位。直接生成语音(直接CSV2S)则受限于多模态复杂性及数据稀缺。为此,本文提出首个统一框架UniCUE,无需中间文本即可直接从手语视频生成语音。核心创新在于融合理解任务(CSR)提供的细粒度视觉-语义线索以指导语音生成。具体包括:姿态感知视觉处理器、实现精确视觉-语义映射的语义对齐池,以及连接理解与生成任务的VisioPhonetic适配器。为支持该框架,构建了大型普通话手语数据集UniCUE-HI,包含14名手势者共11282段视频,涵盖听障与正常听力个体。大量实验表明,UniCUE在多个评估指标上达到当前最优性能。
原文摘要 · Abstract (English)
Cued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Recognition (CSR), which transcribes video content into text. Consequently, a common solution for CSV2S is to integrate CSR with a text-to-speech (TTS) system. However, this pipeline relies on text as an intermediate medium, which may lead to error propagation and temporal misalignment between speech and CS video dynamics. In contrast, directly generating audio speech from CS video (direct CSV2S) often suffers from the inherent multimodal complexity and the limited availability of CS data. To address these challenges, we propose UniCUE, the first unified framework for CSV2S that directly generates speech from CS videos without relying on intermediate text. The core innovation of UniCUE lies in integrating an understanding task (CSR) that provides fine-grained CS visual-semantic cues to guide speech generation. Specifically, UniCUE incorporates a pose-aware visual processor, a semantic alignment pool that enables precise visual-semantic mapping, and a VisioPhonetic adapter to bridge the understanding and generation tasks within a unified architecture. To support this framework, we construct UniCUE-HI, a large-scale Mandarin CS dataset containing 11282 videos from 14 cuers, including both hearing-impaired and normal-hearing individuals. Extensive experiments on this dataset demonstrate that UniCUE achieves state-of-the-art performance across multiple evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。