针对非洲黑人社会设计情感识别模型,提升对话AI的准确性与文化适应性。
Evaluation of Conversational Agents: Understanding Culture, Context and Environment in Emotion Detection
- 融合语音与图像数据,用三层CNN和新AFME算法识别七种基本情绪及讽刺。
- 在非洲黑人语境下实现85%至96%的情感识别准确率。
- 关注文化差异,助力对话AI在多元社会中的伦理合规与可信部署。
当前面部生物识别、社交媒体图像标记及人机交互等应用依赖于高效的情绪识别技术。然而,这些应用在真实场景中的表现受限于边缘案例和文化背景的考量。尽管已有大量通用方法尝试模拟人类情感(包括讽刺),但地理与文化差异的影响仍未充分探索,而这对解决伦理问题和提升对话AI性能至关重要。本文聚焦于非洲黑人社会中对话AI的应用挑战,提出一种结合语音与图像数据的情感预测模型,实现85%至96%的准确率。该模型采用三层卷积神经网络,引入新型音频帧均值表达(AFME)算法,并强化预处理与后处理阶段。最终方案提升了对话AI在复杂文化环境下的情绪识别可信度。
原文摘要 · Abstract (English)
Valuable decisions and highly prioritized analysis now depend on applications such as facial biometrics, social media photo tagging, and human robots interactions. However, the ability to successfully deploy such applications is based on their efficiencies on tested use cases taking into consideration possible edge cases. Over the years, lots of generalized solutions have been implemented to mimic human emotions including sarcasm. However, factors such as geographical location or cultural difference have not been explored fully amidst its relevance in resolving ethical issues and improving conversational AI (Artificial Intelligence). In this paper, we seek to address the potential challenges in the usage of conversational AI within Black African society. We develop an emotion prediction model with accuracies ranging between 85% and 96%. Our model combines both speech and image data to detect the seven basic emotions with a focus on also identifying sarcasm. It uses 3-layers of the Convolutional Neural Network in addition to a new Audio-Frame Mean Expression (AFME) algorithm and focuses on model pre-processing and post-processing stages. In the end, our proposed solution contributes to maintaining the credibility of an emotion recognition system in conversational AIs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。