arXiv:2510.18036cs.SDcs.LG2025-10被引 1

在超低功耗设备上实现语音文本联合情绪识别

Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware

  • 采用量化变压器与冻结关键词嵌入的晚期融合架构
  • 在1.8MB内存内实现实时推理,延迟21-23ms,F1提升6.3%
  • 专为Edge TPU设计,适合可穿戴设备等边缘场景

将情绪识别系统部署于小型、低功耗且需隐私保护的真实环境仍面临重大挑战,尤其适用于紧张监测、冲突缓解和响应式可穿戴设备等场景,此时云端方案不切实际。多模态情绪识别虽因深度学习取得进展,但多数系统仍难以在超受限边缘设备上部署。以往工作通常依赖高性能硬件,缺乏实时性或仅使用单模态输入。本文提出一种面向边缘TPU的硬件感知情绪识别系统,通过晚融合架构结合声学与语言特征。设计整合了量化变压器声学模型与来自DSResNet-SE网络的冻结关键词嵌入,实现1.8MB内存占用下实时推理,延迟21-23ms。采用MicroFrontend与MLTK确保训练与部署中频谱图对齐。在通过Coral Dev Board Micro麦克风采集的重录分段IEMOCAP样本上,相较单模态基线,宏F1提升6.3%。结果表明,通过任务特定融合与硬件引导设计,可在微控制器级边缘平台实现准确、实时的多模态情绪推断。

原文摘要 · Abstract (English)

Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict de-escalation, and responsive wearables, where cloud-based solutions are impractical. Multimodal emotion recognition has advanced through deep learning, but most systems remain unsuitable for deployment on ultra-constrained edge devices. Prior work typically relies on powerful hardware, lacks real-time performance, or uses unimodal input. This paper addresses that gap by presenting a hardware-aware emotion recognition system that combines acoustic and linguistic features using a late-fusion architecture optimised for Edge TPU. The design integrates a quantised transformer-based acoustic model with frozen keyword embeddings from a DSResNet-SE network, enabling real-time inference within a 1.8MB memory budget and 21-23ms latency. The pipeline ensures spectrogram alignment between training and deployment using MicroFrontend and MLTK. Evaluation on re-recorded, segmented IEMOCAP samples captured through the Coral Dev Board Micro microphone shows a 6.3% macro F1 improvement over unimodal baselines. This work demonstrates that accurate, real-time multimodal emotion inference is achievable on microcontroller-class edge platforms through task-specific fusion and hardware-guided model design.

情绪识别边缘计算多模态Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。