构建多模态犬情绪数据集,揭示标注者特征与信息模式对情绪判断的影响。
CREMD: Crowd-Sourced Emotional Multimodal Dogs Dataset
- 采集923段视频,分三类呈现方式测试视觉/音频上下文对情绪识别的影响。
- 无音频时加上下文显著提升标注一致性,音频增强标注信心但结果不明确。
- 非养犬者和男性标注者一致性强于养犬者和女性,专业人士表现最优。
犬情绪识别在改善人犬互动、兽医护理及自动化犬只福祉监测系统中具有重要意义。然而,由于情绪判断主观性强且缺乏标准化真值方法,准确解读犬情绪仍具挑战。本文提出CREMD(Crowd-sourced Emotional Multimodal Dogs Dataset),一个涵盖多种呈现模式(如上下文、音频、视频)及标注者特征(如是否养犬、性别、专业经验)的综合性数据集。数据集包含923个视频片段,以三种模式呈现:无上下文或音频、有上下文无音频、有上下文及音频。我们分析了来自不同背景参与者(包括养犬者、专业人士及不同人口统计群体)的标注,识别影响可靠情绪识别的因素。研究发现:(1) 添加视觉上下文显著提升标注一致性;音频线索影响不明确,因设计缺陷(缺少无上下文带音频条件)及清洁音频资源有限;(2) 与预期相反,非养犬者和男性标注者的一致性高于养犬者和女性;专业人士一致性最高,符合预期;(3) 音频存在显著提高标注者对愤怒与恐惧等特定情绪的信心。
原文摘要 · Abstract (English)
Dog emotion recognition plays a crucial role in enhancing human-animal interactions, veterinary care, and the development of automated systems for monitoring canine well-being. However, accurately interpreting dog emotions is challenging due to the subjective nature of emotional assessments and the absence of standardized ground truth methods. We present the CREMD (Crowd-sourced Emotional Multimodal Dogs Dataset), a comprehensive dataset exploring how different presentation modes (e.g., context, audio, video) and annotator characteristics (e.g., dog ownership, gender, professional experience) influence the perception and labeling of dog emotions. The dataset consists of 923 video clips presented in three distinct modes: without context or audio, with context but no audio, and with both context and audio. We analyze annotations from diverse participants, including dog owners, professionals, and individuals with varying demographic backgrounds and experience levels, to identify factors that influence reliable dog emotion recognition. Our findings reveal several key insights: (1) while adding visual context significantly improved annotation agreement, our findings regarding audio cues are inconclusive due to design limitations (specifically, the absence of a no-context-with-audio condition and limited clean audio availability); (2) contrary to expectations, non-owners and male annotators showed higher agreement levels than dog owners and female annotators, respectively, while professionals showed higher agreement levels, aligned with our initial hypothesis; and (3) the presence of audio substantially increased annotators' confidence in identifying specific emotions, particularly anger and fear.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。