arXiv:2606.00670cs.SDcs.AI2026-06

上脸表情能提升嘈杂环境下语音识别的鲁棒性,不靠字词信息,而是帮系统判断信心。

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

论文配图:Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty
图 1 · 摘自论文原文
  • 对比四组线索:仅声、加嘴部、加上脸、全脸,测试不同组合在噪声中的表现
  • 0dB信噪比下,加嘴部特征使准确率提升7.94%,加上脸特征改善模型校准度
  • 上脸情绪线索虽不直接提高识别率,但能增强系统对噪声的适应能力和信心判断

面对面语音理解本质上是多模态的,融合声音信号与口部动作、面部表情、头部运动等社交相关线索。现有音视频语音系统通常以口部为首要视觉信息源,而将面部表情视为独立的情绪识别任务。本文探究在声音质量下降时,上脸情绪线索是否有助于音视频句子识别,超越声音和口部线索。基于CREMA-D音视频情感语音数据集,我们训练了四种线索条件下的特征分类器:仅声音(A)、声音+口部/下脸特征(A+M)、声音+上脸特征(A+U)、声音+口部与上脸特征(A+M+U)。在干净音频及粉红噪声条件下,分别在+10 dB、+5 dB、0 dB信噪比下,使用演员无关划分进行评估。结果表明,口部/下脸特征在声音退化时带来显著鲁棒性提升:0 dB SNR下,A+M相比A准确率提高0.0794,95%置信区间[0.0296, 0.1298]。上脸情绪线索作用更复杂:尽管A+M+U相对于A+M的直接准确率提升较小,但全脸模型在各信噪比下始终具有更好校准性,并在噪声条件下优于随机打乱的上脸控制组。这表明,情绪面部信息可能在声学不确定下支持多模态鲁棒性和置信度估计,而非直接编码词汇内容。研究强调了社交表达性面部线索在以人为本的音视频交互系统中的潜在价值。

原文摘要 · Abstract (English)

Face-to-face speech comprehension is inherently multimodal, integrating acoustic signals with visible articulation, facial expression, head motion, and other socially relevant cues. While audiovisual speech systems typically focus on the mouth region as the primary visual source of linguistic information, affective facial expressions are often treated separately as emotion-recognition targets. This paper investigates whether upper-face affective information contributes to audiovisual sentence recognition beyond audio and mouth-region cues, particularly under acoustic degradation. Using the CREMA-D audiovisual emotional speech corpus, we train feature-based sentence classifiers under four cue conditions: audio only (A), audio plus mouth/lower-face features (A+M), audio plus upper-face features (A+U), and audio plus both mouth and upper-face features (A+M+U). Models are evaluated on clean audio and pink-noise conditions at +10 dB, +5 dB, and 0 dB SNR using actor-independent splits. Results show that mouth/lower-face features provide substantial robustness benefits under degraded audio. At 0 dB SNR, A+M improves accuracy over A by 0.0794, with an actor-bootstrap 95% confidence interval of [0.0296, 0.1298]. Upper-face affective cues exhibit a more nuanced effect. Although the direct accuracy gain of A+M+U over A+M is small, full-face models consistently improve calibration across SNR levels and outperform shuffled upper-face controls under noisy conditions. These findings suggest that affective facial information may support multimodal robustness and confidence estimation under acoustic uncertainty without directly encoding lexical content. More broadly, the study highlights the potential role of socially expressive facial cues in human-centered audiovisual interaction systems.

音视频识别情绪线索鲁棒性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。