通过多维注意力机制提升唇读模型在复杂场景下的识别准确率
MA-LipNet: Multi-Dimensional Attention Networks for Robust Lipreading
- 设计三重注意力模块,分别从通道、时空维度净化视觉特征
- 在CMLR和GRID数据集上将字符错误率降低至15.2%和9.8%
- 适合需要高鲁棒性唇读技术的安防与语音辅助场景
唇读技术通过分析无声视频中的口型运动来解码语音内容,在公共安全等领域具有重要应用价值。然而,由于发音动作细微,现有方法普遍存在特征区分度低、泛化能力差的问题。为此,本文从时序、空间和通道三个维度对视觉特征进行净化,提出一种名为多维注意力唇读网络(MA-LipNet)的新方法。其核心在于依次引入三个专用注意力模块:首先使用通道注意力(CA)自适应重校准通道特征,削弱不相关信息干扰;随后采用两种不同粒度的时空注意力模块——联合时空注意力(JSTA)通过统一权重图实现粗粒度滤波,分离时空注意力(SSTA)则分别建模时空注意力以实现细粒度优化。在CMLR和GRID数据集上的大量实验表明,MA-LipNet显著降低了字符错误率(CER)和词错误率(WER),验证了其有效性与先进性。本工作强调了多维特征精炼对鲁棒视觉语音识别的重要性。
原文摘要 · Abstract (English)
Lipreading, the technology of decoding spoken content from silent videos of lip movements, holds significant application value in fields such as public security. However, due to the subtle nature of articulatory gestures, existing lipreading methods often suffer from limited feature discriminability and poor generalization capabilities. To address these challenges, this paper delves into the purification of visual features from temporal, spatial, and channel dimensions. We propose a novel method named Multi-Attention Lipreading Network(MA-LipNet). The core of MA-LipNet lies in its sequential application of three dedicated attention modules. Firstly, a \textit{Channel Attention (CA)} module is employed to adaptively recalibrate channel-wise features, thereby mitigating interference from less informative channels. Subsequently, two spatio-temporal attention modules with distinct granularities-\textit{Joint Spatial-Temporal Attention (JSTA)} and \textit{Separate Spatial-Temporal Attention (SSTA)}-are leveraged to suppress the influence of irrelevant pixels and video frames. The JSTA module performs a coarse-grained filtering by computing a unified weight map across the spatio-temporal dimensions, while the SSTA module conducts a more fine-grained refinement by separately modeling temporal and spatial attentions. Extensive experiments conducted on the CMLR and GRID datasets demonstrate that MA-LipNet significantly reduces the Character Error Rate (CER) and Word Error Rate (WER), validating its effectiveness and superiority over several state-of-the-art methods. Our work highlights the importance of multi-dimensional feature refinement for robust visual speech recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。