融合图像与传感器数据,提升缺基础设施地区的空气质量预测精度。
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
- 用轻量CNN提取图像和污染数据特征,通过语义对齐融合。
- 在印尼8000+样本上达92.02%准确率,RMSE仅7.70。
- 适合边缘设备部署,部分传感器失效时仍稳定预测。
资源匮乏地区因传感器分布稀疏、设施有限,空气污染监测面临挑战。本文提出AQFusionNet,一种多模态深度学习框架,用于鲁棒的空气质量指数(AQI)预测。该框架结合地面大气图像与污染物浓度数据,采用轻量级CNN主干网络(MobileNetV2、ResNet18、EfficientNet-B0),通过语义对齐的嵌入空间融合视觉与传感器特征,实现高精度高效预测。在印度和尼泊尔超过8000个样本上的实验表明,AQFusionNet持续优于单模态基线模型,使用EfficientNet-B0主干时分类准确率达92.02%,均方根误差(RMSE)为7.70。相比单模态方法,性能提升达18.5%,同时保持低计算开销,适合在边缘设备部署。该模型为基础设施有限环境下的AQI监测提供了可扩展、实用的解决方案,即使在部分传感器不可用时也具备强鲁棒性预测能力。
原文摘要 · Abstract (English)
Air pollution monitoring in resource-constrained regions remains challenging due to sparse sensor deployment and limited infrastructure. This work introduces AQFusionNet, a multimodal deep learning framework for robust Air Quality Index (AQI) prediction. The framework integrates ground-level atmospheric imagery with pollutant concentration data using lightweight CNN backbones (MobileNetV2, ResNet18, EfficientNet-B0). Visual and sensor features are combined through semantically aligned embedding spaces, enabling accurate and efficient prediction. Experiments on more than 8,000 samples from India and Nepal demonstrate that AQFusionNet consistently outperforms unimodal baselines, achieving up to 92.02% classification accuracy and an RMSE of 7.70 with the EfficientNet-B0 backbone. The model delivers an 18.5% improvement over single-modality approaches while maintaining low computational overhead, making it suitable for deployment on edge devices. AQFusionNet provides a scalable and practical solution for AQI monitoring in infrastructure-limited environments, offering robust predictive capability even under partial sensor availability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。