融合摄像头与气象数据,预测富士山景观可见度。
FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility
- 用图像分类概率和天气数据晚期融合预测可见度。
- 当天预测准确率约89%,次日预测达84%。
- 适合做环境预测与多模态学习研究的人参考。
自然地标如富士山的可见度是旅游规划和游客体验的关键因素,但受快速变化的气象条件影响,难以预测。我们提出FujiView,一个用于预测景观可见度的多模态学习框架与数据集,通过融合网络摄像头图像与结构化气象数据实现预测。采用晚期融合策略,将图像提取的类别概率与数值天气特征结合,将可见度分为五个等级。当前数据集包含超过10万张富士山周边40多个摄像头拍摄的同步及预报气象数据,将持续扩展;该数据集将公开以支持环境预测研究。实验表明,基于YOLO的视觉特征在短时预测(如“即时预报”和“同日视域”)中占主导地位,而天气驱动的预报在超过+1天的时间尺度上逐渐成为主要预测信号。晚期融合始终取得最高准确率,在同日预测中达到约0.89的精度,次日预测最高达84%。这些结果使景观可见度预测(SVF)成为多模态学习的新基准任务。
原文摘要 · Abstract (English)
Visibility of natural landmarks such as Mount Fuji is a defining factor in both tourism planning and visitor experience, yet it remains difficult to predict due to rapidly changing atmospheric conditions. We present FujiView, a multimodal learning framework and dataset for predicting scenic visibility by fusing webcam imagery with structured meteorological data. Our late-fusion approach combines image-derived class probabilities with numerical weather features to classify visibility into five categories. The dataset currently comprises over 100,000 webcam images paired with concurrent and forecasted weather conditions from more than 40 cameras around Mount Fuji, and continues to expand; it will be released to support further research in environmental forecasting. Experiments show that YOLO-based vision features dominate short-term horizons such as "nowcasting" and "samedaycasting", while weather-driven forecasts increasingly take over as the primary predictive signal beyond $+1$d. Late fusion consistently yields the highest overall accuracy, achieving ACC of approx 0.89 for same-day prediction and up to 84% for next-day forecasts. These results position Scenic Visibility Forecasting (SVF) as a new benchmark task for multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。