用手机声波识别表情,无需摄像头也不怕隐私泄露。
Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
- 通过手机扬声器与面部轮廓的超声回波分析表情特征。
- 在20人实验中实现90%以上准确率,跨用户表现优于现有方法10%。
- 无需麦克风阵列,小样本下也能泛化,适合日常场景应用。
表情识别在内容推荐和心理健康等领域具有广泛应用前景,传统方法依赖摄像头或可穿戴设备,存在隐私问题且增加设备负担。现有基于声音的方法在训练与推理数据分布不一致时性能下降明显。本文提出FacER+,一种主动式声学面部表情识别系统,无需外部麦克风阵列。该系统利用智能手机扬声器发射近超声信号,通过分析其在3D面部轮廓与耳塞之间的回波,提取表情特征,有效抑制背景噪声,并支持不同用户间低样本下的表达识别。我们设计了一种对比外部注意力模型,以在跨用户场景中持续学习表达特征,减少分布差异。在包含20名志愿者、有无口罩的多样化真实场景下进行的大量实验表明,FacER+能以超过90%的准确率识别六类常见面部表情,性能优于当前领先的声学感知方法10%。FacER+为面部表情识别提供了一种鲁棒且实用的解决方案。
原文摘要 · Abstract (English)
Expression recognition holds great promise for applications such as content recommendation and mental healthcare by accurately detecting users' emotional states. Traditional methods often rely on cameras or wearable sensors, which raise privacy concerns and add extra device burdens. In addition, existing acoustic-based methods struggle to maintain satisfactory performance when there is a distribution shift between the training dataset and the inference dataset. In this paper, we introduce FacER+, an active acoustic facial expression recognition system, which eliminates the requirement for external microphone arrays. FacER+ extracts facial expression features by analyzing the echoes of near-ultrasound signals emitted between the 3D facial contour and the earpiece speaker on a smartphone. This approach not only reduces background noise but also enables the identification of different expressions from various users with minimal training data. We develop a contrastive external attention-based model to consistently learn expression features across different users, reducing the distribution differences. Extensive experiments involving 20 volunteers, both with and without masks, demonstrate that FacER+ can accurately recognize six common facial expressions with over 90% accuracy in diverse, user-independent real-life scenarios, surpassing the performance of the leading acoustic sensing methods by 10%. FacER+ offers a robust and practical solution for facial expression recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。