用三个轻量模块实现眼神、表情和说话人识别,适合资源受限的辅助设备。
Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
- 分模块设计:用CNN、深度CNN和LSTM分别处理眼神、表情和语音
- 三类任务准确率超93%,最高达97.8%(表情识别)
- 模型轻量且可独立部署,适合嵌入式辅助设备
开发全面的辅助技术需要视觉与听觉感知的无缝融合。本研究评估了一种受'Smart Eye'等感知系统启发的模块化架构的可行性。我们提出并基准测试了三个独立的感知模块:用于眼态检测(困倦/专注)的卷积神经网络(CNN),用于面部表情识别的深度CNN,以及用于语音说话人识别的长短期记忆网络(LSTM)。基于Eyes Image、FER2013和自定义音频数据集,各模型分别达到93.0%、97.8%和96.89%的准确率。研究表明,轻量级、领域专用模型可在离散任务上实现高保真度,为未来在资源受限的辅助设备中实现实时多模态集成奠定了验证基础。
原文摘要 · Abstract (English)
Developing comprehensive assistive technologies requires the seamless integration of visual and auditory perception. This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye.' We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification. Utilizing the Eyes Image, FER2013, and customized audio datasets, our models achieved accuracies of 93.0%, 97.8%, and 96.89%, respectively. This study demonstrates that lightweight, domain-specific models can achieve high fidelity on discrete tasks, establishing a validated foundation for future real-time, multimodal integration in resource-constrained assistive devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。