用哈达玛乘积融合手工特征与深度特征,提升真实微笑识别准确率。
HadaSmileNet: Hadamard fusion of handcrafted and deep-learning features for enhancing facial emotion recognition of genuine smiles
- 通过哈达玛乘积实现深度特征与生理标记特征的直接融合。
- 在四个数据集上达到最高准确率,最高提升5.0个百分点。
- 计算效率高,参数减少26%,适合实时情感计算应用。
真实与伪装情绪的区分是模式识别中的基本挑战,对社会科学研究、医疗健康和人机交互有重要意义。尽管多任务学习框架结合深度模型与手工D-Marker特征在微笑情绪识别中表现良好,但存在辅助任务监督和复杂损失平衡带来的计算低效问题。本文提出HadaSmileNet,一种通过无参数乘法交互直接融合基于Transformer的表示与生理学驱动的D-Marker特征的新颖特征融合框架。通过系统评估15种融合策略,证明哈达玛乘法融合在保持计算高效的同时实现最优性能。该方法在四个基准数据集上建立新SOTA:UvA-NEMO(88.7%,+0.8)、MMI(99.7%)、SPOS(98.5%,+0.7)和BBC(100%,+5.0)。综合计算分析显示,相比多任务方案,参数量减少26%,训练更简化;特征可视化表明,通过直接引入领域知识增强了判别能力。该框架兼具高效性与有效性,特别适合需要实时情感计算的多媒体数据挖掘场景。
原文摘要 · Abstract (English)
The distinction between genuine and posed emotions represents a fundamental pattern recognition challenge with significant implications for data mining applications in social sciences, healthcare, and human-computer interaction. While recent multi-task learning frameworks have shown promise in combining deep learning architectures with handcrafted D-Marker features for smile facial emotion recognition, these approaches exhibit computational inefficiencies due to auxiliary task supervision and complex loss balancing requirements. This paper introduces HadaSmileNet, a novel feature fusion framework that directly integrates transformer-based representations with physiologically grounded D-Markers through parameter-free multiplicative interactions. Through systematic evaluation of 15 fusion strategies, we demonstrate that Hadamard multiplicative fusion achieves optimal performance by enabling direct feature interactions while maintaining computational efficiency. The proposed approach establishes new state-of-the-art results for deep learning methods across four benchmark datasets: UvA-NEMO (88.7 percent, +0.8), MMI (99.7 percent), SPOS (98.5 percent, +0.7), and BBC (100 percent, +5.0). Comprehensive computational analysis reveals 26 percent parameter reduction and simplified training compared to multi-task alternatives, while feature visualization demonstrates enhanced discriminative power through direct domain knowledge integration. The framework's efficiency and effectiveness make it particularly suitable for practical deployment in multimedia data mining applications that require real-time affective computing capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。