arXiv:2501.04799cs.CL2025-01被引 1

用预训练模型生成手语动作,让听障者更清晰地理解口语。

Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model

  • 基于音视频文生语音模型,从文本生成手部与唇部动作。
  • 在音素级别识别准确率达约77%。
  • 适合听障语言辅助与多模态交互研究者。

本文提出一种自动生成有声手语(ACSG)的新方法,这是一种为听障人士设计的视觉语言表达系统,有助于更准确地传达口语。我们探索了迁移学习策略,利用预训练的音视频自回归文生语音模型(AVTacotron2),将其重构为根据文本输入推断有声手语手势和口型动作的模型。在两个公开数据集上进行了实验,其中一个为本研究专门采集。通过自动有声手语识别系统评估性能,结果表明,在音素级别上的解码准确率约为77%,验证了该方法的有效性。

原文摘要 · Abstract (English)

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach.

有声手语音视频模型听障辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。