用连续情绪空间分析宠物叫声,更精准识别情绪强弱与类型。
Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
- 构建二维情绪空间,用连续值替代传统分类标签。
- 在42,553条语音上训练,情绪强度预测相关系数达0.90(正向)和0.72(唤醒)。
- 多任务学习提升特征表达,适合宠物情感交互与临床诊断应用。
传统宠物情绪识别依赖离散分类,难以处理情绪模糊性和强度差异。本文提出一种连续的效价-唤醒(Valence-Arousal, VA)模型,将情绪表示为二维空间中的连续值。通过自动标注算法,大规模标注了42,553条宠物叫声样本。采用多任务学习框架,联合训练VA回归与辅助任务(情绪、体型、性别),以增强特征学习。所提出的Audio Transformer模型在验证集上取得效价Pearson相关系数r = 0.9024,唤醒r = 0.7155,有效区分如“领地性”与“愉悦”等易混淆类别。本研究首次建立宠物叫声分析的连续VA框架,为人类-宠物互动、兽医诊断及行为训练提供更丰富的情绪表达,具备向消费级AI宠物情绪翻译产品部署的潜力。
原文摘要 · Abstract (English)
Traditional pet emotion recognition from vocalizations, based on discrete classification, struggles with ambiguity and capturing intensity variations. We propose a continuous Valence-Arousal (VA) model that represents emotions in a two-dimensional space. Our method uses an automatic VA label generation algorithm, enabling large-scale annotation of 42,553 pet vocalization samples. A multi-task learning framework jointly trains VA regression with auxiliary tasks (emotion, body size, gender) to enhance prediction by improving feature learning. Our Audio Transformer model achieves a validation Valence Pearson correlation of r = 0.9024 and an Arousal r = 0.7155, effectively resolving confusion between discrete categories like "territorial" and "happy." This work introduces the first continuous VA framework for pet vocalization analysis, offering a more expressive representation for human-pet interaction, veterinary diagnostics, and behavioral training. The approach shows strong potential for deployment in consumer products like AI pet emotion translators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。