arXiv:2509.14430eess.AScs.SD2025-09

通过多通道差异输入提升智能眼镜语音识别鲁棒性

Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses

  • 融合波束成形、麦克风选择与轻量侧话检测,构建互补输入机制
  • 在真实和模拟数据集上实现最高18.0%的词错误率相对降低
  • 适合需要高鲁棒性语音识别的可穿戴设备开发者

随着智能眼镜等可穿戴设备在人工智能助手中的广泛应用,佩戴者语音识别(WSR)正成为下一代人机交互的关键。然而,在实际环境中,旁侧对话语音的干扰仍是WSR的重大挑战,可能导致下游自然语言处理任务出现累积误差。本文提出一种新型多通道差异自动语音识别(ASR)方法,用于提升智能眼镜上的稳健WSR性能。该系统从多个互补的前端获取差异输入,包括波束成形、麦克风选择以及轻量级侧话检测模型。在模拟和真实数据集上的评估表明,所提系统优于传统方法,词错误率最高降低18.0%。

原文摘要 · Abstract (English)

With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate.

语音识别可穿戴设备抗干扰智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。