通过差异学习与跨视图交互提升唇读模型对左右唇部细微差别的捕捉能力
RAL:Redundancy-Aware Lipreading Model Based on Differential Learning with Symmetric Views
- 设计对称视图的差异学习策略,挖掘左右唇部非对称信息
- 在LRW和LRW-1000上分别提升4.3%和3.8%准确率
- 适合关注细粒度视觉语义的唇读研究者
唇读旨在通过分析唇部运动序列解读说话人言语。当前多数模型将左右唇部分视为对称整体,未充分挖掘其差异。然而,左右唇部并不总是对称,其细微差别蕴含丰富语义信息。本文提出基于对称视图的差异学习策略(DLSV)以解决此问题。此外,输入图像常包含与识别无关的冗余信息,会降低模型性能。为此,我们设计冗余感知操作(RAO)以减少冗余。最后,为利用对称视图间及视图内部的关联信息,进一步提出自适应跨视图交互模块(ACVI)。在LRW与LRW-1000数据集上的实验充分验证了方法的有效性。
原文摘要 · Abstract (English)
Lip reading involves interpreting a speaker's speech by analyzing sequences of lip movements. Currently, most models regard the left and right halves of the lips as a symmetrical whole, lacking a thorough investigation of their differences. However, the left and right halves of the lips are not always symmetrical, and the subtle differences between them contain rich semantic information. In this paper, we propose a differential learning strategy with symmetric views (DLSV) to address this issue. Additionally, input images often contain a lot of redundant information unrelated to recognition results, which can degrade the model's performance. We present a redundancy-aware operation (RAO) to reduce it. Finally, to leverage the relational information between symmetric views and within each view, we further design an adaptive cross-view interaction module (ACVI). Experiments on LRW and LRW-1000 datasets fully demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。