arXiv:2511.15131eess.AScs.CL2025-11中稿 · ICASSP 2026被引 4

构建首个大规模真人标注音频定位数据集,提升真实场景下音频检索性能。

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

  • 基于真人标注构建1009条训练音频的大规模数据集
  • 在真实数据上微调模型,召回率提升10.4点
  • 适合音频定位、语音理解等研究者使用

我们提出CASTELLA,一个用于音频时刻检索(AMR)任务的人工标注音频基准数据集。尽管AMR具有广泛应用潜力,但此前缺乏基于真实世界数据的标准化基准。早期研究仅在合成数据集上训练模型,且评估依赖于少于100个样本的标注数据,导致性能报告不可靠。为提升真实环境下的应用效果,我们构建了大规模人工标注的CASTELLA数据集,包含1009、213和640条音频样本,分别用于训练、验证和测试,规模是此前数据集的24倍。我们还基于CASTELLA建立了基线模型。实验表明,在合成数据预训练后,于CASTELLA上微调的模型在[email protected]上比仅在合成数据上训练的模型高出10.4分。CASTELLA已公开发布于https://h-munakata.github.io/CASTELLA-demo/。

原文摘要 · Abstract (English)

We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The initial study of AMR trained the models solely on synthetic datasets. Moreover, the evaluation is based on an annotated dataset of fewer than 100 samples. This resulted in less reliable reported performance. To ensure performance for applications in real-world environments, we present CASTELLA, a large-scale manually annotated AMR dataset. CASTELLA consists of 1009, 213, and 640 audio recordings for training, validation, and test splits, respectively, which is 24 times larger than the previous dataset. We also establish a baseline model for AMR using CASTELLA. Our experiments demonstrate that a model fine-tuned on CASTELLA after pre-training on the synthetic data outperformed a model trained solely on the synthetic data by 10.4 points in [email protected]. CASTELLA is publicly available in https://h-munakata.github.io/CASTELLA-demo/.

音频检索数据集语音理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。