arXiv:2511.02495cs.CVcs.CL2025-11NeurIPS被引 2

构建首个大规模火情多模态数据集,助力智能防火研究。

DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding

  • 整合2.5万段真实火灾视频与22.5万张高清图像,覆盖多样场景
  • 含边界框与详细文本描述,支持检测、生成与推理任务
  • 专为火灾理解设计,适合安全系统与视觉语言模型研究者

多模态模型在图像生成和推理任务中表现优异,但其在火灾领域的应用受限于缺乏高质量标注的公开数据集。为此,我们提出DetectiumFire,一个大规模多模态数据集,包含22.5k张高分辨率火灾相关图像和2.5k段真实火灾视频,涵盖多种火情类型、环境与风险等级。数据经传统计算机视觉标签(如边界框)和详尽文本提示双重标注,支持合成数据生成与火灾风险推理等应用。相比现有基准,DetectiumFire在规模、多样性与数据质量上均有显著提升,有效减少冗余并增强真实场景覆盖。我们在目标检测、基于扩散的图像生成及视觉-语言推理等多个任务中验证了其有效性,证明该数据集能推动火灾相关研究,助力智能安全系统开发。数据集已开放获取,欢迎社区使用。

原文摘要 · Abstract (English)

Recent advances in multi-modal models have demonstrated strong performance in tasks such as image generation and reasoning. However, applying these models to the fire domain remains challenging due to the lack of publicly available datasets with high-quality fire domain annotations. To address this gap, we introduce DetectiumFire, a large-scale, multi-modal dataset comprising of 22.5k high-resolution fire-related images and 2.5k real-world fire-related videos covering a wide range of fire types, environments, and risk levels. The data are annotated with both traditional computer vision labels (e.g., bounding boxes) and detailed textual prompts describing the scene, enabling applications such as synthetic data generation and fire risk reasoning. DetectiumFire offers clear advantages over existing benchmarks in scale, diversity, and data quality, significantly reducing redundancy and enhancing coverage of real-world scenarios. We validate the utility of DetectiumFire across multiple tasks, including object detection, diffusion-based image generation, and vision-language reasoning. Our results highlight the potential of this dataset to advance fire-related research and support the development of intelligent safety systems. We release DetectiumFire to promote broader exploration of fire understanding in the AI community. The dataset is available at https://kaggle.com/datasets/38b79c344bdfc55d1eed3d22fbaa9c31fad45e27edbbe9e3c529d6e5c4f93890

多模态火灾识别数据集视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。