VSR模型靠语言统计而非视觉感知,无法真正像人一样读唇。
The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

- 用多层级指标对比模型与人类在真实数据上的表现
- 模型准确率高但错误模式与人类不同,依赖训练词频而非视觉信息
- 适合关注模型可解释性与真实感知能力的研究者
视觉语音识别(VSR)模型在标准评测中已超越人类唇读能力,但这是否意味着具备类人视觉语音感知?我们通过MaFI词级唇读数据集,使用词、字符、音素和视觉音素级别指标,比较三种VSR系统与人类基线的表现。尽管模型整体准确率更高,其成功与失败的词汇与人类不同。仅基于文本n-gram、给定少量初始音素的基线模型已接近人类表现。VSR的词级错误更受训练词频影响,而非词汇的视觉可辨识度。视觉音素准确率、混淆矩阵及人类-模型相关性分析进一步表明,模型在人类最难识别的视觉音素上表现最好,且对视觉清晰度依赖极弱。研究揭示:VSR系统主要依赖训练数据中的语言线索,而非将视觉特征整合为有意义的词汇,未能实现真正的视觉语义绑定。
原文摘要 · Abstract (English)
Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare three VSR systems with human baselines on the MaFI word-level lipreading dataset using word, character, phoneme, and viseme-level metrics. Although models achieve higher overall accuracy, they succeed and fail on different words than humans. A text-only n-gram baseline given only a few initial phonemes rivals human lipreading. VSR word-level errors are consistently better explained by training word frequency than by the visual informativeness of words. Viseme accuracies, confusion matrices and human-model correlations further show that models gain most on visemes humans find hardest, and show much weaker dependence on visual clarity. Our work demonstrates that VSR systems rely primarily on language cues from training data rather than visual perception, failing to bind visual features into meaningful words.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。