arXiv:2606.22399cs.SD2026-06

构建首个带呼号标注的空管语音数据集,支持精准指令识别与语音检索。

ATCCaps: A Call-Sign-Aware Speech Dataset for Air Traffic Control Recognition

论文配图:ATCCaps: A Call-Sign-Aware Speech Dataset for Air Traffic Control Recognition
图 1 · 摘自论文原文
  • 基于真实空管录音,结合ADS-B数据与大模型生成,实现呼号级语音-文本对齐。
  • 含202.94小时音频、17万条语句、922个唯一呼号,支持呼号匹配与检索评估。
  • 适合空管语音识别、呼号解析、多模态模型训练等安全关键场景研究者使用。

呼号是空管通信中至关重要的安全实体,用于标识每条语音指令的目标飞机。本文提出ATCCaps,一个具备呼号感知能力的空管语音数据集,采用逐字幕级别的音视频-文本监督。该数据集基于真实空管无线电信号录音构建,包含202.94小时经筛选的音频、170,385条语句和922个唯一归一化呼号。其构建流程融合置信度感知的转录解析、基于ADS-B的呼号元数据、呼号归一化、规则式质量过滤及大语言模型辅助的字幕生成。每个保留样本均配有转录描述、呼号描述和空管风格字幕,支持自动语音识别(ASR)评估、呼号匹配与呼号感知的音视频-文本检索。我们通过划分统计、呼号覆盖率、已见/未见呼号分析、过滤审计及字幕质量评估对数据集进行了全面表征。评估子集源自人工标注的ATCO2-test-set,可提供基于人工转录的基准评估。结果表明,ATCCaps提供了可扩展的、以音频为依据的呼号监督机制,而字幕分析凸显了显式验证呼号与数字准确性的必要性。参考的ASR与CLAP基线模型展示了该数据集在呼号感知空管语音建模中的可用性。

原文摘要 · Abstract (English)

Call signs are safety-critical entities in air traffic control (ATC) communications because they identify the target aircraft of each spoken instruction. This paper presents ATCCaps, a call-sign-aware ATC speech dataset with caption-level audio-text supervision. Built from real ATC radiotelephony recordings, ATCCaps contains 202.94 hours of curated audio, 170,385 utterances, and 922 unique normalized call signs. The construction pipeline combines confidence-aware transcript parsing, ADS-B-derived call-sign metadata, call-sign normalization, rule-based quality filtering, and LLM-assisted caption generation. Each retained sample is paired with transcript descriptions, call-sign descriptions, and ATC-style captions, supporting ASR evaluation, call-sign matching, and call-sign-aware audio-text retrieval. We further characterize ATCCaps through split statistics, call-sign coverage, seen/unseen call-sign analysis, filtering audits, and caption quality evaluation. The evaluation subset is derived from the human-annotated ATCO2-test-set, enabling reference evaluation with manual transcripts. Results show that ATCCaps provides scalable audio-grounded call-sign supervision, while caption analysis highlights the need to explicitly validate call-sign and numeric fidelity. Reference ASR and CLAP-based baselines demonstrate the usability of ATCCaps for call-sign-aware ATC speech modeling.

语音识别空管系统呼号识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。