用街景视频和语言模型判断灾后建筑是否可住,提升救援效率。
Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
- 通过语言引导将街景视频对齐到建筑,识别入口堵塞等可解释特征
- 两阶段方法在飓风后评估中达F1 0.848,优于单阶段基线
- 输出中间结果便于定位错误,适合应急管理和地理信息系统使用
灾后建筑入住情况对优先级划分、检查、设施恢复及资源公平分配至关重要。航拍影像覆盖快但常遗漏立面与通行线索,街景影像虽能捕捉细节却稀疏且难与地块对齐。我们提出FacadeTrack,一种基于语言引导的街景级框架,可将全景视频关联至地块,校正视角至立面,并提取可解释属性(如入口堵塞、临时遮盖、局部废墟),支持两种决策策略:透明的一阶段规则与感知与保守推理分离的两阶段设计。在两次飓风海琳灾后调查中,两阶段方法达到精确率0.927、召回率0.781、F1分数0.848,优于单阶段基线(精确率0.943,召回率0.728,F1 0.822)。除精度外,中间属性与空间诊断揭示了残余误差位置与原因,支持针对性质量控制。该流程提供可审计、可扩展的入住评估,适用于地理空间与应急管理体系集成。
原文摘要 · Abstract (English)
Building-level occupancy after disasters is vital for triage, inspections, utility re-energization, and equitable resource allocation. Overhead imagery provides rapid coverage but often misses facade and access cues that determine habitability, while street-view imagery captures those details but is sparse and difficult to align with parcels. We present FacadeTrack, a street-level, language-guided framework that links panoramic video to parcels, rectifies views to facades, and elicits interpretable attributes (for example, entry blockage, temporary coverings, localized debris) that drive two decision strategies: a transparent one-stage rule and a two-stage design that separates perception from conservative reasoning. Evaluated across two post-Hurricane Helene surveys, the two-stage approach achieves a precision of 0.927, a recall of 0.781, and an F-1 score of 0.848, compared with the one-stage baseline at a precision of 0.943, a recall of 0.728, and an F-1 score of 0.822. Beyond accuracy, intermediate attributes and spatial diagnostics reveal where and why residual errors occur, enabling targeted quality control. The pipeline provides auditable, scalable occupancy assessments suitable for integration into geospatial and emergency-management workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。