arXiv:2410.00349cs.RO2024-10被引 7

用3D人脸模型生成数据,提升机器人互动中情绪预测准确率

Data Augmentation for 3DMM-based Arousal-Valence Prediction for HRI

  • 基于3DMM生成缺失情绪数据,增强训练样本多样性
  • 在SEWA数据集上实现0.793的情绪预测相关系数
  • 适合做人机交互中实时情绪识别的开发者参考

人类通过多种交流渠道互动,如肢体动作或面部表情传递意图。这推动了情绪预测模型的发展,特别是从面部表情预测唤醒度(arousal)与效价(valence)。然而,在人机交互(HRI)场景下,要实现高精度预测仍具挑战,需应对多主体、复杂条件和丰富表情变化。本文提出一种基于3D可变形模型(3DMM)的数据增强方法,用于提升AV预测器性能。该方法针对SEWA数据集中未充分覆盖的AV区域生成合成序列,该数据集是目前最全面的连续AV标注数据集。实验表明,该方法显著提升了实时应用中的预测准确性和鲁棒性。模型在SEWA上的预测相关系数达到0.793(arousal和valence)。

原文摘要 · Abstract (English)

Humans use multiple communication channels to interact with each other. For instance, body gestures or facial expressions are commonly used to convey an intent. The use of such non-verbal cues has motivated the development of prediction models. One such approach is predicting arousal and valence (AV) from facial expressions. However, making these models accurate for human-robot interaction (HRI) settings is challenging as it requires handling multiple subjects, challenging conditions, and a wide range of facial expressions. In this paper, we propose a data augmentation (DA) technique to improve the performance of AV predictors using 3D morphable models (3DMM). We then utilize this approach in an HRI setting with a mediator robot and a group of three humans. Our augmentation method creates synthetic sequences for underrepresented values in the AV space of the SEWA dataset, which is the most comprehensive dataset with continuous AV labels. Results show that using our DA method improves the accuracy and robustness of AV prediction in real-time applications. The accuracy of our models on the SEWA dataset is 0.793 for arousal and valence.

情绪识别3DMM数据增强人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。