arXiv:2607.08236cs.CV2026-07

提升事件相机唇读性能,通过轨迹感知与发音单元引导的时序聚合。

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

论文配图:TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
图 1 · 摘自论文原文
  • 先局部时序建模再自适应空间聚合,保留细微唇部运动轨迹。
  • 引入发音单元引导的时序模块,词级识别准确率显著提升。
  • 结合教师-学生一致性训练,增强对噪声事件流的鲁棒性。

事件相机唇读近年成为视觉语音识别的新方向,得益于事件相机的高时间分辨率和运动敏感性。然而,现有方法通常在充分时序建模前进行空间压缩,可能抑制对区分相似唇动至关重要的稀疏局部运动轨迹。此外,多数方法主要在词分类层面优化时序表示,导致发音结构约束不足。为此,我们提出一种增强时序建模的事件相机唇读框架:首先引入轨迹感知差异聚合(TDA),在每个空间位置进行局部时序建模后,再进行自适应空间聚合;其次提出发音单元引导聚合(VGA),由CTC解码器与发音单元引导门控聚合分支组成,注入发音单元级序列监督,提升最终时序聚合效果;最后采用EMA教师-学生训练策略,增强在强事件扰动下的鲁棒性。在DVS-Lip基准上的实验验证了设计的有效性,消融实验进一步证明TDA、VGA及师生一致性的作用。定性结果表明,基于CTC的时序建模能从事件流中学习有意义的发音单元感知结构。

原文摘要 · Abstract (English)

Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typically perform spatial compression before sufficient temporal modeling, which may suppress sparse and localized motion trajectories that are crucial for distinguishing similar lip movements. Moreover, most current approaches optimize temporal representations mainly at the word-classification level, leaving the underlying articulatory structure weakly constrained. To address these limitations, we propose a temporally enhanced framework for event-based lip reading. First, we introduce Trajectory-Aware Differential Aggregation (TDA), which performs local temporal modeling at each spatial location before adaptive spatial aggregation. Second, we propose Viseme-Guided Aggregation (VGA), a unified temporal module composed of a CTC decoder and a viseme-guided gated aggregation branch, which injects viseme-aware sequence supervision and improves final temporal aggregation for word recognition. Third, we incorporate an EMA teacher--student training strategy to enhance robustness under strong event perturbations. Experiments on the DVS-Lip benchmark verify the effectiveness of the proposed design, and extensive ablation studies further validate the contributions of TDA, VGA, and teacher--student consistency. Qualitative decoding results also demonstrate that the proposed CTC-based temporal modeling learns meaningful viseme-aware structure from event streams.

唇读事件相机时序建模发音单元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。