针对警用对讲机通信的语音识别难题,提出改进方案并公开数据集。
Speech Recognition for Analysis of Police Radio Communication
- 用62,000条人工转录数据微调模型,提升识别准确率
- 定制模型性能接近人工水平,远超通用语音模型
- 适合研究执法通信分析、低资源语音识别方向的学者
全球警察部门广泛使用双向对讲机进行协调。这些警用广播通信(BPC)是日常警务活动与应急响应的独特信息来源,但尚未被转录,其自然语音特性也使自动语音识别(ASR)极具挑战。本文收集了约62,000条手动转录的无线电传输记录(总计约46小时音频),评估现代识别模型在该领域的可行性。我们测试了现成语音识别器、在BPC数据上微调的模型以及定制端到端模型。结果表明,人类与机器转录在此领域均具挑战性:大型通用ASR模型表现不佳,但经微调后模型可达到接近人类的性能水平。本研究为未来工作指明方向,包括短语识别与警用沟通中的潜在误解分析。我们公开了语料库及标注流程,以支持该领域进一步研究。
原文摘要 · Abstract (English)
Police departments around the world use two-way radio for coordination. These broadcast police communications (BPC) are a unique source of information about everyday police activity and emergency response. Yet BPC are not transcribed, and their naturalistic audio properties make automatic transcription challenging. We collect a corpus of roughly 62,000 manually transcribed radio transmissions (~46 hours of audio) to evaluate the feasibility of automatic speech recognition (ASR) using modern recognition models. We evaluate the performance of off-the-shelf speech recognizers, models fine-tuned on BPC data, and customized end-to-end models. We find that both human and machine transcription is challenging in this domain. Large off-the-shelf ASR models perform poorly, but fine-tuned models can reach the approximate range of human performance. Our work suggests directions for future work, including analysis of short utterances and potential miscommunication in police radio interactions. We make our corpus and data annotation pipeline available to other researchers, to enable further research on recognition and analysis of police communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。