arXiv:2603.17108cs.CV2026-03

用社交媒体图片实时估测厘米级洪水深度,助力交通韧性提升

LLM-Powered Flood Depth Estimation from Social Media Imagery: A Vision-Language Model Framework with Mechanistic Interpretability for Transportation Resilience

  • 基于抖音数据构建百万级合成图像集,微调视觉语言模型实现街景洪水深度估计
  • 平均误差低于0.97厘米,深水场景准确率超96.8%,浅水更依赖简单提示
  • 揭示关键特征层并压缩参数量76%-80%,支持实时部署与抗遮挡鲁棒性

城市洪涝威胁交通网络连续性,但现有系统无法提供厘米级、实时的街景洪水深度信息,难以支撑动态路径规划、电动车安全及自动驾驶需求。本文提出FloodLlama,一种基于开源视觉语言模型(VLM)的连续洪水深度估计算法,依托抖音数据构建的多模态感知管道。通过约19万张合成图像训练,覆盖7种车类、4种天气及41个深度层级(0-40厘米,1厘米分辨率)。采用渐进式课程学习实现粗到细的训练策略,使用QLoRA对LLaMA 3.2-11B Vision进行微调。34,797次评估显示:浅水场景下简单提示效果更优,深水则需链式思维推理提升性能。模型在深水场景平均绝对误差低于0.97厘米,5厘米内准确率超过93.7%,浅水超过96.8%。五阶段机制可解释框架识别出第23层为关键深度编码过渡层,支持选择性微调,使可训练参数减少76%-80%且保持精度。三级配置在真实数据上达98.62%准确率,并具备强抗遮挡能力。基于抖音的数据管道在底特律676帧标注样本中验证可行性,表明该方案可实现无基础设施、可扩展的实时洪水感知,直接推动电动车安全、自动驾驶落地和交通韧性管理。

原文摘要 · Abstract (English)

Urban flooding poses an escalating threat to transportation network continuity, yet no operational system currently provides real-time, street-level flood depth information at the centimeter resolution required for dynamic routing, electric vehicle (EV) safety, and autonomous vehicle (AV) operations. This study presents FloodLlama, a fine-tuned open-source vision-language model (VLM) for continuous flood depth estimation from single street-level images, supported by a multimodal sensing pipeline using TikTok data. A synthetic dataset of approximately 190000 images was generated, covering seven vehicle types, four weather conditions, and 41 depth levels (0-40 cm at 1 cm resolution). Progressive curriculum training enabled coarse-to-fine learning, while LLaMA 3.2-11B Vision was fine-tuned using QLoRA. Evaluation across 34797 trials reveals a depth-dependent prompt effect: simple prompts perform better for shallow flooding, whereas chain-of-thought (CoT) reasoning improves performance at greater depths. FloodLlama achieves a mean absolute error (MAE) below 0.97 cm and Acc@5cm above 93.7% for deep flooding, exceeding 96.8% for shallow depths. A five-phase mechanistic interpretability framework identifies layer L23 as the critical depth-encoding transition and enables selective fine-tuning that reduces trainable parameters by 76-80% while maintaining accuracy. The Tier 3 configuration achieves 98.62% accuracy on real-world data and shows strong robustness under visual occlusion. A TikTok-based data pipeline, validated on 676 annotated flood frames from Detroit, demonstrates the feasibility of real-time, crowd-sourced flood sensing. The proposed framework provides a scalable, infrastructure-free solution with direct implications for EV safety, AV deployment, and resilient transportation management.

洪水估计视觉语言模型交通韧性实时感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。