首个工业巡检安全多模态基准数据集,助力智能巡检系统可靠评估。
Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios
- 基于真实机器人在5类工业场景采集,覆盖2239个巡检点位
- 每条实例含像素级分割与安全等级标签,支持细粒度感知
- 融合7种传感模态,适用于多模态安全分析与模型训练
随着工业智能化和无人巡检的快速发展,复杂动态工业环境中人工智能系统的可靠感知与安全评估已成为预测性维护与自主巡检部署的关键瓶颈。现有公开数据集普遍存在仿真数据、单模态感知或缺乏细粒度物体级标注的问题,难以支撑工业基础模型的鲁棒场景理解与多模态安全推理。为此,我们发布InspecSafe-V1,首个面向工业巡检安全评估的多模态基准数据集,源自真实巡检机器人在实际环境中的常规作业。该数据集涵盖隧道、电力设施、烧结设备、油气石化厂及煤输送栈桥共5类典型工业场景,由41台轮式与轨道式巡检机器人在2,239个有效巡检站点采集,共生成5,013个巡检实例。每个实例提供可见光图像中关键物体的像素级分割标注,并依据实际巡检任务给出语义场景描述与对应的安全等级标签。同时,包含红外视频、音频、深度点云、雷达点云、气体浓度、温度与湿度等七种同步传感模态,支持多模态异常识别、跨模态融合与综合安全评估。
原文摘要 · Abstract (English)
With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in complex and dynamic industrial sites has become a key bottleneck for deploying predictive maintenance and autonomous inspection. Most public datasets remain limited by simulated data sources, single-modality sensing, or the absence of fine-grained object-level annotations, which prevents robust scene understanding and multimodal safety reasoning for industrial foundation models. To address these limitations, InspecSafe-V1 is released as the first multimodal benchmark dataset for industrial inspection safety assessment that is collected from routine operations of real inspection robots in real-world environments. InspecSafe-V1 covers five representative industrial scenarios, including tunnels, power facilities, sintering equipment, oil and gas petrochemical plants, and coal conveyor trestles. The dataset is constructed from 41 wheeled and rail-mounted inspection robots operating at 2,239 valid inspection sites, yielding 5,013 inspection instances. For each instance, pixel-level segmentation annotations are provided for key objects in visible-spectrum images. In addition, a semantic scene description and a corresponding safety level label are provided according to practical inspection tasks. Seven synchronized sensing modalities are further included, including infrared video, audio, depth point clouds, radar point clouds, gas measurements, temperature, and humidity, to support multimodal anomaly recognition, cross-modal fusion, and comprehensive safety assessment in industrial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。