arXiv:2506.07572cs.CVcs.CL2025-06

让唇读模型忽略说话人差异,提升跨人识别准确率。

Learning Speaker-Invariant Visual Features for Lipreading

  • 用隐式与显式双重解耦机制分离说话人特有特征。
  • 在多个公开数据集上超越现有最佳方法,泛化能力显著提升。
  • 适合需要跨说话人唇读的场景,如语音辅助识别系统。

唇读是一项将视觉唇部运动转化为口语文本的挑战性跨模态任务。现有方法提取的视觉特征常包含说话人特异属性(如形状、颜色、纹理),导致视觉与文本间产生虚假相关性,影响识别准确率并限制模型泛化能力。为此,本文提出SIFLip框架,通过两个互补的解耦模块(隐式解耦与显式解耦)实现说话人无关的视觉特征学习。由于不同说话人在念相同词汇时唇动与语音具有语义一致性,隐式解耦模块利用稳定的文本嵌入作为监督信号,隐式学习跨说话人的共性视觉表征;同时,在主唇读流程中设计说话人识别子任务,通过梯度反转显式过滤并解耦个性化视觉特征。实验表明,SIFLip在多个公开数据集上显著提升泛化性能,优于当前最优方法。

原文摘要 · Abstract (English)

Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape, color, texture), which introduce spurious correlations between vision and text. These correlations lead to suboptimal lipreading accuracy and restrict model generalization. To address this challenge, we introduce SIFLip, a speaker-invariant visual feature learning framework that disentangles speaker-specific attributes using two complementary disentanglement modules (Implicit Disentanglement and Explicit Disentanglement) to improve generalization. Specifically, since different speakers exhibit semantic consistency between lip movements and phonetic text when pronouncing the same words, our implicit disentanglement module leverages stable text embeddings as supervisory signals to learn common visual representations across speakers, implicitly decoupling speaker-specific features. Additionally, we design a speaker recognition sub-task within the main lipreading pipeline to filter speaker-specific features, then further explicitly disentangle these personalized visual features from the backbone network via gradient reversal. Experimental results demonstrate that SIFLip significantly enhances generalization performance across multiple public datasets. Experimental results demonstrate that SIFLip significantly improves generalization performance across multiple public datasets, outperforming state-of-the-art methods.

唇读特征解耦跨人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。