arXiv:2608.04072cs.SDcs.HC2026-08

用深度学习实时识别重叠声音顺序,提升人机交互中的听觉感知能力。

Deep Learning for Real-Time Sound Order Recognition in Human-Robot Interaction

  • 多分支CNN融合梅尔谱、梅尔倒谱系数和短时傅里叶变换特征
  • 在相同音量下准确率达99%,变音量下仍保持91%准确率
  • 适合需要实时声音顺序判断的机器人交互场景

在人机交互中,识别重叠声音的时间顺序是一个未充分探索的挑战,与救援人员检测等应用直接相关。本文提出一种基于深度学习的实时声音顺序识别框架,使用可记录的蜂鸣器发出不同非语言声音(猫叫、狗吠、直升机声)。采用多分支卷积神经网络处理梅尔谱图、梅尔频率倒谱系数(MFCC)和短时傅里叶变换(STFT)特征,并通过注意力机制融合关键时间线索。在相同振幅、变振幅及未见声音条件下进行实验,系统在平衡重叠下达到99%准确率,变振幅下为91%,在未见数据上经归一化后达74%。结果表明,深度学习可可靠识别重叠条件下的声音顺序,支持实际人机交互场景。尽管实验基于受控合成重叠,研究还报告了延迟基准,证明其实时可行性,并讨论泛化性、生态效度及真实环境部署挑战。

原文摘要 · Abstract (English)

Recognizing the temporal order of overlapping sounds is an underexplored challenge in human-robot interaction (HRI), with direct relevance to applications such as first responder detection systems. This paper presents a deep learning framework for real-time sound order recognition using recordable buzzers that emit distinct non-verbal sounds (cat meows, dog barks, helicopter noises). A multi-branch convolutional neural network (CNN) processes Mel spectrograms, Mel-frequency cepstral coefficients (MFCCs), and short-time Fourier transform (STFT) features, with an attention-based fusion mechanism to emphasize critical temporal cues. Experiments were conducted under same-amplitude, varied-amplitude, and unseen sound conditions. The proposed system achieved 99% accuracy in balanced overlaps, 91% under amplitude variation, and 74% on unseen test data with normalization. These results demonstrate that deep learning can reliably recognize sound order in overlapping conditions, supporting practical HRI scenarios. While experiments were conducted on carefully controlled synthetic overlaps, we additionally report latency benchmarks demonstrating real-time feasibility and provide an extended discussion on generalization, ecological validity, and deployment challenges in real-room environments.

声音识别人机交互实时系统深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。