用多模态数据提升空管呼叫信号识别在极端情况下的鲁棒性。
Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding
- 融合语音与文本等多模态信息,重建受损通话记录中的呼叫信号。
- 在极端场景下性能提升最高达15%,且模型参数更少、微调更快。
- 适合对可靠性要求高的空管语音系统研发人员使用。
基于机器学习的辅助系统需在各种场景下保持稳健,尤其在空中交通管制(ATC)领域更为关键。系统的鲁棒性在边缘案例中尤为突出,例如高词错误率(WER)的噪声录音或因截断导致的部分转录。为增强呼叫信号识别与理解(CRU)在边缘情况下的鲁棒性,我们提出多模态呼叫信号-指令恢复模型(CCR)。该架构使边缘案例性能最高提升15%。我们在第二个提出的架构CallSBERT上验证了这一效果:该模型参数更少,可显著加快微调速度,并在微调过程中更具鲁棒性,优于当前最先进的CRU模型。此外,优化边缘案例表现能显著提升整个运行范围内的准确率。
原文摘要 · Abstract (English)
Operational machine-learning based assistant systems must be robust in a wide range of scenarios. This hold especially true for the air-traffic control (ATC) domain. The robustness of an architecture is particularly evident in edge cases, such as high word error rate (WER) transcripts resulting from noisy ATC recordings or partial transcripts due to clipped recordings. To increase the edge-case robustness of call-sign recognition and understanding (CRU), a core tasks in ATC speech processing, we propose the multimodal call-sign-command recovery model (CCR). The CCR architecture leads to an increase in the edge case performance of up to 15%. We demonstrate this on our second proposed architecture, CallSBERT. A CRU model that has less parameters, can be fine-tuned noticeably faster and is more robust during fine-tuning than the state of the art for CRU. Furthermore, we demonstrate that optimizing for edge cases leads to a significantly higher accuracy across a wide operational range.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。