仅用音频实现三类情绪分类,适合低算力设备部署。
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
- 用10个SVM和6个神经网络堆叠集成,提升分类性能。
- 真实电影数据集上达到86%准确率,模拟数据67%。
- 专为轻量级设备设计,适合家庭影音系统应用。
情绪识别在提升人机交互中至关重要,尤其在电影推荐系统中理解情感内容尤为关键。尽管融合音频与视频的多模态方法已证明有效,但其对高性能图形计算的依赖限制了在个人电脑或家庭音视频系统等资源受限设备上的部署。为此,本研究提出一种新颖的纯音频集成学习框架,可将电影场景分为三类情绪:好、中性、差。该模型在堆叠集成架构中融合10个支持向量机与6个神经网络以增强分类表现。设计了定制化的数据预处理流程,包括特征提取、异常值处理与特征工程,以优化音频输入中的情感信息。在模拟数据集上实验获得67%准确率,在来自15部不同影片的真实数据集上取得86%的显著准确率。结果表明,基于音频的轻量级情绪识别方法具有广泛消费级应用潜力,兼具计算效率与强分类能力。
原文摘要 · Abstract (English)
Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have demonstrated effectiveness, their reliance on high-performance graphical computing limits deployment on resource-constrained devices such as personal computers or home audiovisual systems. To address this limitation, this study proposes a novel audio-only ensemble learning framework capable of classifying movie scenes into three emotional categories: Good, Neutral, and Bad. The model integrates ten support vector machines and six neural networks within a stacking ensemble architecture to enhance classification performance. A tailored data preprocessing pipeline, including feature extraction, outlier handling, and feature engineering, is designed to optimize emotional information from audio inputs. Experiments on a simulated dataset achieve 67% accuracy, while a real-world dataset collected from 15 diverse films yields an impressive 86% accuracy. These results underscore the potential of audio-based, lightweight emotion recognition methods for broader consumer-level applications, offering both computational efficiency and robust classification capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。