arXiv:2412.17907cs.HCcs.CL2024-12被引 6

融合多模态信息,实现更客观的情绪识别

A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language

  • 整合面部表情、语音、语言和身体动作四类信号
  • 在模拟场景中验证了情绪判断的可靠性提升
  • 适合临床心理评估与治疗辅助场景

传统心理评估依赖人工观察与解读,易受主观性、偏见、疲劳和不一致影响。为此,本文提出一种多模态情绪识别系统,为心理学家、精神科医生和临床医师提供标准化、客观且数据驱动的辅助工具。系统融合面部表情、语音、口语内容及身体运动分析,捕捉人类评估中常被忽略的细微情绪线索。通过多模态协同,提升情绪状态评估的鲁棒性与全面性,降低误诊与过度诊断风险。初步在模拟真实环境条件下测试,证明该系统具备提供可靠情绪洞察的潜力,有望成为传统心理评估的有效补充,适用于临床与治疗场景。

原文摘要 · Abstract (English)

Traditional psychological evaluations rely heavily on human observation and interpretation, which are prone to subjectivity, bias, fatigue, and inconsistency. To address these limitations, this work presents a multimodal emotion recognition system that provides a standardised, objective, and data-driven tool to support evaluators, such as psychologists, psychiatrists, and clinicians. The system integrates recognition of facial expressions, speech, spoken language, and body movement analysis to capture subtle emotional cues that are often overlooked in human evaluations. By combining these modalities, the system provides more robust and comprehensive emotional state assessment, reducing the risk of mis- and overdiagnosis. Preliminary testing in a simulated real-world condition demonstrates the system's potential to provide reliable emotional insights to improve the diagnostic accuracy. This work highlights the promise of automated multimodal analysis as a valuable complement to traditional psychological evaluation practices, with applications in clinical and therapeutic settings.

情绪识别多模态临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。