首个面向工业港口环境的弱标签音频数据集,助力真实场景下的声音事件检测研究。
Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring

- 基于两个固定传感器采集22小时工业港口音频,覆盖26类典型声事件。
- 采用弱标签策略,提供两种标注版本以应对标注者差异,支持鲁棒模型训练。
- 适用于工业安全监控、边缘设备实时音频分析等实际应用场景。
Soroll-IA 是在西班牙瓦伦西亚真实工业港口环境中,通过两个固定传感节点采集的弱标签环境音频数据集。数据集包含约22小时音频,分割为7,396个音频片段,涵盖26种代表工业港口声学活动的声事件类别,如起重机警报、火车运行、交通噪声及其他物流与工业声响。录音条件极为复杂,存在强背景噪声、远距离声源和频繁事件重叠。所有音频片段由领域专家按弱标签策略标注,仅标明片段中是否存在声事件,不提供时间定位。为应对标注者间差异,发布两个真值版本:一是非交叉验证版本(至少一名专家标注即视为存在),二是更保守的交叉验证版本(需至少三分之二专家一致)。该数据集旨在支持音频标记、弱监督声事件检测及真实工业声学条件下机器学习研究。基准测试采用两种互补架构:CNN14(代表高容量卷积模型)和MobileNetV2(适合低资源边缘设备实时分类)。据当前所知,Soroll-IA是唯一专注于工业港口声学环境的数据集,致力于推动安全关键与运营监控场景下的鲁棒环境声音分析。数据集已在线公开,采用CC BY-NC 4.0许可。
原文摘要 · Abstract (English)
Soroll-IA is a weakly labeled environmental audio dataset recorded in a real-world industrial port environment in Valencia (Spain) using two fixed sensing nodes. The dataset comprises approximately 22 hours of audio segmented into 7,396 clips and covers 26 sound event classes representative of industrial port acoustic activity commonly observed in such environments, such as crane sirens, train movements, traffic, and other logistical and industrial sounds. Recordings were captured under highly challenging acoustic conditions, including strong background noise, long-distance sources, and frequent event overlap. All audio clips were annotated by domain experts following a weak labeling strategy, where tags indicate the presence of sound events within a clip without temporal localization. To account for inter-annotator variability, two ground-truth versions are released: one without cross-validation, where a class is considered present if annotated by at least one expert, and a second, more conservative version based on cross-validation, where agreement by at least two-thirds of the annotators is required. The dataset is intended to support research in audio tagging, weakly supervised sound event detection, and machine learning under realistic industrial acoustic conditions. Benchmark results are provided using two complementary architectures: CNN14 representing high-capacity convolutional models for audio tagging, and MobileNetV2, selected for its suitability in real-time classification on low-resource edge devices. To the best of current knowledge, Soroll-IA constitutes an available dataset dedicated exclusively to industrial port acoustic environments, aiming to foster advances in robust environmental sound analysis for safety-critical and operational monitoring applications. The dataset is available online and collected under Attribution-NonCommercial 4.0 International license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。