arXiv:2410.05986eess.AScs.SD2024-10被引 1

用模拟数据提升智能眼镜语音识别实时性。

The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge

  • 通过多重重叠率模拟数据+一对一匹配训练,减少真实与模拟数据偏差。
  • 融合智能眼镜的IMU数据,使语音识别在实时场景下准确率提升。
  • 适合研究可穿戴设备语音识别、多模态融合的开发者参考。

在双人对话场景中,佩戴智能眼镜的一方若能实时转录并显示说话内容,将为翻译、理解等后续任务提供先验信息。然而,从智能眼镜获取的多模态数据极为稀缺。为此,我们提出利用具有多种重叠率的模拟数据,并采用一对一匹配的训练策略,以缩小模型训练中真实数据与模拟数据之间的偏差。此外,将智能眼镜中的惯性测量单元(IMU)数据引入模型,有助于音频信号实现更优的实时语音识别性能。

原文摘要 · Abstract (English)

In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a priori information for subsequent tasks such as translation and comprehension. Meanwhile, multi-modal data captured from the smart glasses is scarce. Therefore, we propose utilizing simulation data with multiple overlap rates and a one-to-one matching training strategy to narrow down the deviation for the model training between real and simulated data. In addition, combining IMU unit data in the model can assist the audio to achieve better real-time speech recognition performance.

语音识别多模态智能眼镜实时

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。