arXiv:2609.05782cs.AI2026-09

将大模型压缩到设备端,实现安全高效的火灾理解。

Distilling Vision-Language Models for On-Device Fire Understanding

  • 用教师-学生框架蒸馏专用于火灾理解的视觉语言模型。
  • 轻量级模型保留了大部分教师模型的火灾识别能力。
  • 适合资源受限、安全关键场景的智能设备部署。

视觉语言模型(VLMs)通过分析场景语义上下文,可有效减少传统火灾检测系统的误报,但其庞大的模型规模难以在嵌入式火灾传感器上部署。本文研究如何将领域专用的VLM压缩至完全可在设备端运行,同时保持火灾检测所需的安全关键行为。我们提出一种教师-学生知识蒸馏框架,将针对火灾理解微调的大规模VLM压缩为轻量级学生模型。在多个VLM系列和模型规模上的实验表明,紧凑的学生模型保留了教师模型绝大部分的火灾理解能力。我们进一步将蒸馏模型部署于商用Detectium火灾检测传感器,并联合评估推理准确率、延迟与内存占用。结果表明,压缩与部署不仅影响准确率,也改变模型失效模式,其中Qwen2.5-0.5B展现出最佳综合性能。研究成果为资源受限、安全关键场景中领域专用VLM的部署提供了广泛指导。

原文摘要 · Abstract (English)

Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers' fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.

视觉语言模型模型压缩火灾检测边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。