构建火灾场景360视频基准,评估模型在恶劣环境下的感知与记忆能力
Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos
- 构建包含228段360度火灾训练视频的多条件数据集
- 人类在物体变换匹配任务中准确率达83.5%,大模型表现显著落后
- 聚焦烟雾、弱光等极端环境下视觉推理与记忆能力评测
现代AI系统在可靠性至关重要的环境中表现最差——如烟雾弥漫、能见度低、结构变形的场景。每年数万名消防员因情境感知失效而受伤。我们提出Fire360,一个用于评估安全关键型灭火场景下感知与推理能力的基准。该数据集包含228段来自专业训练的360度视频,覆盖低光、热畸变等多种条件,标注了动作片段、物体位置及退化元信息。支持五项任务:视觉问答、时间动作描述、物体定位、安全关键推理和变换物体检索(TOR)。TOR测试模型能否在无配对场景中将原始样本与火损版本匹配,评估其变换不变识别能力。人类专家在TOR上达到83.5%准确率,而GPT-4o等模型表现明显不足,暴露出在退化条件下推理的缺陷。通过发布Fire360及其评估套件,我们旨在推动不仅能‘看’,还能‘记’‘思’‘行’于不确定环境中的模型发展。数据集获取地址:https://uofi.box.com/v/fire360dataset。
原文摘要 · Abstract (English)
Modern AI systems struggle most in environments where reliability is critical - scenes with smoke, poor visibility, and structural deformation. Each year, tens of thousands of firefighters are injured on duty, often due to breakdowns in situational perception. We introduce Fire360, a benchmark for evaluating perception and reasoning in safety-critical firefighting scenarios. The dataset includes 228 360-degree videos from professional training sessions under diverse conditions (e.g., low light, thermal distortion), annotated with action segments, object locations, and degradation metadata. Fire360 supports five tasks: Visual Question Answering, Temporal Action Captioning, Object Localization, Safety-Critical Reasoning, and Transformed Object Retrieval (TOR). TOR tests whether models can match pristine exemplars to fire-damaged counterparts in unpaired scenes, evaluating transformation-invariant recognition. While human experts achieve 83.5% on TOR, models like GPT-4o lag significantly, exposing failures in reasoning under degradation. By releasing Fire360 and its evaluation suite, we aim to advance models that not only see, but also remember, reason, and act under uncertainty. The dataset is available at: https://uofi.box.com/v/fire360dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。