arXiv:2510.07837cs.CVcs.MM2025-10中稿 · AIML-Systems-2025被引 1

直接将手语视频转为语音,无需中间文本,提升沟通效率。

IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries

  • 端到端设计,跳过文本中间步骤,减少延迟与错误积累。
  • 在ASL-Citizen-1500和WLASL-100上达到72.01%与78.67%的准确率。
  • 创新使用非极大值抑制算法,有效识别非语法连续手语序列。

手语到口语音频的转换对连接听力与语言障碍者具有重要意义。本文聚焦于孤立手语序列而非连贯语法手语视频,此类数据适用于教育场景与手语提示接口。为此,我们提出IsoSignVid2Aud,一种无需中间文本表示的端到端框架,可将一系列可能无语法的连续手语视频直接转为语音,实现即时交流并避免多阶段系统中的延迟与级联误差。方法结合I3D特征提取模块、专用特征变换网络与音频生成流水线,并引入新型非极大值抑制(NMS)算法,用于非语法连续序列中手语的时间定位。实验表明,在ASL-Citizen-1500与WLASL-100数据集上,Top-1准确率分别为72.01%和78.67%,音频质量指标PESQ为2.67,STOI为0.73,表明输出语音具备可懂性。代码已公开于:https://github.com/BheeshmSharma/IsoSignVid2Aud_AIMLsystems-2025。

原文摘要 · Abstract (English)

Sign language to spoken language audio translation is important to connect the hearing- and speech-challenged humans with others. We consider sign language videos with isolated sign sequences rather than continuous grammatical signing. Such videos are useful in educational applications and sign prompt interfaces. Towards this, we propose IsoSignVid2Aud, a novel end-to-end framework that translates sign language videos with a sequence of possibly non-grammatic continuous signs to speech without requiring intermediate text representation, providing immediate communication benefits while avoiding the latency and cascading errors inherent in multi-stage translation systems. Our approach combines an I3D-based feature extraction module with a specialized feature transformation network and an audio generation pipeline, utilizing a novel Non-Maximal Suppression (NMS) algorithm for the temporal detection of signs in non-grammatic continuous sequences. Experimental results demonstrate competitive performance on ASL-Citizen-1500 and WLASL-100 datasets with Top-1 accuracies of 72.01\% and 78.67\%, respectively, and audio quality metrics (PESQ: 2.67, STOI: 0.73) indicating intelligible speech output. Code is available at: https://github.com/BheeshmSharma/IsoSignVid2Aud_AIMLsystems-2025.

手语翻译端到端语音生成AI助残

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。