首个统一框架,让手语、口型和语音协同生成文本。
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
- 设计通用架构处理手语、口型、语音多模态输入
- 在四项任务上表现媲美或超越专用模型
- 发现口型作为非手动线索能显著提升手语识别
语音是人类沟通的主要方式,推动了自动语音识别(ASR)技术的发展。然而,以语音为中心的系统天然排除了聋人或听力障碍者。视觉替代方案如手语和唇读提供了有效途径,近年来手语翻译(SLT)和视觉语音识别(VSR)的进步提升了无音频沟通能力。但这些模态长期被孤立研究,其整合在统一框架中的探索仍不足。本文提出首个可处理手语、口型与音频多种组合的统一框架,用于生成口语文本。核心目标包括:(i) 设计能有效处理异构输入的通用、模态无关架构;(ii) 探索模态间未充分研究的协同效应,尤其关注口型在手语理解中的非手动线索作用;(iii) 实现与各专项任务顶尖模型相当甚至更优的性能。基于该框架,在SLT、VSR、ASR及音视频语音识别任务上均达到或超越现有最先进水平。分析揭示:将口型显式建模为独立模态,可显著提升手语翻译性能,捕捉关键非手动信息。
原文摘要 · Abstract (English)
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual alternatives such as sign language and lip reading offer effective substitutes, and recent advances in Sign Language Translation (SLT) and Visual Speech Recognition (VSR) have improved audio-less communication. Yet, these modalities have largely been studied in isolation, and their integration within a unified framework remains underexplored. In this paper, we propose the first unified framework capable of handling diverse combinations of sign language, lip movements, and audio for spoken-language text generation. We focus on three main objectives: (i) designing a unified, modality-agnostic architecture capable of effectively processing heterogeneous inputs; (ii) exploring the underexamined synergy among modalities, particularly the role of lip movements as non-manual cues in sign language comprehension; and (iii) achieving performance on par with or superior to state-of-the-art models specialized for individual tasks. Building on this framework, we achieve performance on par with or better than task-specific state-of-the-art models across SLT, VSR, ASR, and Audio-Visual Speech Recognition. Furthermore, our analysis reveals a key linguistic insight: explicitly modeling lip movements as a distinct modality significantly improves SLT performance by capturing critical non-manual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。