arXiv:2606.00751cs.CV2026-06

通过显式引入头部姿态信息,提升非正对视角下的唇语识别准确率。

Head-Pose-Aware Visual Speech Recognition with FiLM Modulation

  • 用姿态条件化的FiLM模块动态调节视觉特征,增强对头部转动的鲁棒性。
  • 在LRS2和LRS3数据集上分别达到25.0%和33.2%的词错误率,无需额外数据。
  • 对大偏航角(>30°)样本增益显著,计算开销小,适合真实场景部署。

唇语识别(VSR)旨在从嘴唇运动等视觉线索中识别语音,但其性能受音素模糊性和姿态引起的几何畸变与遮挡的根本限制。现有方法主要依赖语言上下文或隐式不变性,导致非正面视角下视觉表示不够稳健。本文提出一种姿态感知的音素级框架HP-VSR-ResFiLM,显式将头部姿态信息融入视觉特征提取。该框架采用两阶段流程:第一阶段使用姿态条件化的残差特征逐点线性调制(FiLM)块,在2D CNN前端后自适应优化视觉表示;第二阶段利用预训练的NLLB语言模型进行音素到文本重构。在LRS2和LRS3数据集上的实验表明,该方法在相似训练条件下取得具有竞争力的表现,词错误率(WER)分别为25.0%和33.2%,且不依赖额外训练数据。消融实验显示,单个残差FiLM块可稳定降低整体WER,而第3、4层的深层调制对偏航角大于30°的样本提升更明显,同时不影响小角度变化的表现。结果表明,显式的姿态感知特征调制是提升开放环境中VSR鲁棒性的有效且高效方案。

原文摘要 · Abstract (English)

Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme ambiguity and pose-induced variations that introduce geometric distortions and occlusions. Existing approaches mainly rely on linguistic context or implicit invariance, leaving visual representations insufficiently robust under non-frontal views. In this work, we propose a pose-aware phoneme-level framework, termed HP-VSR-ResFiLM, that explicitly incorporates head-pose information into visual feature extraction. The proposed framework adopts a two-stage pipeline consisting of a pose-conditioned visual encoder in Stage 1 and a pretrained NLLB language model in Stage 2 for phoneme-to-text reconstruction. Specifically, Stage 1 incorporates a pose-conditioned residual Feature-wise Linear Modulation (FiLM) block after the 2D CNN frontend to adaptively refine visual representations using head-pose information. Experiments on LRS2 and LRS3 demonstrate that HP-VSR-ResFiLM achieves competitive performance under comparable training conditions, attaining word error rates (WER) of 25.0% and 33.2%, respectively, without relying on additional training data. Ablation studies further show that a single residual FiLM block consistently improves overall WER, while deeper modulation at Layers 3 and 4 provides larger gains for samples with yaw angles greater than 30° without degrading performance for smaller pose variations. These findings demonstrate that explicit pose-aware feature modulation offers an effective and computationally efficient solution for improving VSR robustness in unconstrained settings.

唇语识别姿态感知FiLM视觉语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。