arXiv:2509.04772cs.CVcs.AI2025-09被引 4

用视觉语言模型和知识图谱实现城市洪水深度零样本估算

FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph

  • 结合大模型语义理解与真实物体尺寸知识图谱,避免幻觉
  • 在110张众包图像上达到8.17厘米平均误差,比基线提升20.5%
  • 适合智慧城市建设中实时洪水监测与公众上报应用

及时准确的洪水深度估计对道路通行和应急响应至关重要。尽管近期计算机视觉方法已实现洪水检测,但因依赖固定目标检测器和特定任务训练,存在精度不足和泛化能力差的问题。本文提出FloodVision,一种零样本框架,融合基础视觉语言模型GPT-4o与结构化领域知识图谱。该知识图谱编码常见城市物体(如车辆、行人、基础设施)的典型物理尺寸,使模型推理基于真实世界。FloodVision动态识别图像中可见参照物,从知识图谱中检索已验证高度以减少幻觉,估计淹没比例,并通过统计异常值过滤得出最终深度。在来自MyCoast纽约的110张众包图像上评估,平均绝对误差为8.17厘米,较GPT-4o基线降低10.28厘米,提升20.5%,优于以往基于CNN的方法。系统在不同场景下泛化良好,近实时运行,适用于未来集成至数字孪生平台与市民报告应用,增强智慧城市抗洪韧性。

原文摘要 · Abstract (English)

Timely and accurate floodwater depth estimation is critical for road accessibility and emergency response. While recent computer vision methods have enabled flood detection, they suffer from both accuracy limitations and poor generalization due to dependence on fixed object detectors and task-specific training. To enable accurate depth estimation that can generalize across diverse flood scenarios, this paper presents FloodVision, a zero-shot framework that combines the semantic reasoning abilities of the foundation vision-language model GPT-4o with a structured domain knowledge graph. The knowledge graph encodes canonical real-world dimensions for common urban objects including vehicles, people, and infrastructure elements to ground the model's reasoning in physical reality. FloodVision dynamically identifies visible reference objects in RGB images, retrieves verified heights from the knowledge graph to mitigate hallucination, estimates submergence ratios, and applies statistical outlier filtering to compute final depth values. Evaluated on 110 crowdsourced images from MyCoast New York, FloodVision achieves a mean absolute error of 8.17 cm, reducing the GPT-4o baseline 10.28 cm by 20.5% and surpassing prior CNN-based methods. The system generalizes well across varying scenes and operates in near real-time, making it suitable for future integration into digital twin platforms and citizen-reporting apps for smart city flood resilience.

洪水估计视觉语言模型知识图谱智能城市

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。