用视觉语言模型融合图像与人类知识,生成高质量灾害损毁数据。
Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
- 结合图像与人类认知,利用视觉语言模型生成损毁数据。
- 生成数据提升对建筑、道路等不同损毁程度的分类准确率。
- 适合灾害响应、遥感标注及数据匮乏场景使用。
在人道主义援助与灾害响应(HADR)中,及时准确评估损毁至关重要。当前深度学习方法因类别不平衡、中等损毁样本稀缺以及人工像素标注误差等问题,难以有效泛化。为克服这些局限,本文提出利用先进的视觉语言模型(VLMs)融合图像与人类知识理解,高效生成多样化的基于图像的损毁数据。初步实验结果表明,生成数据质量令人鼓舞,显著提升了对建筑物、道路及基础设施等不同损毁程度场景的分类能力。
原文摘要 · Abstract (English)
It is of crucial importance to assess damages promptly and accurately in humanitarian assistance and disaster response (HADR). Current deep learning approaches struggle to generalize effectively due to the imbalance of data classes, scarcity of moderate damage examples, and human inaccuracy in pixel labeling during HADR situations. To accommodate for these limitations and exploit state-of-the-art techniques in vision-language models (VLMs) to fuse imagery with human knowledge understanding, there is an opportunity to generate a diversified set of image-based damage data effectively. Our initial experimental results suggest encouraging data generation quality, which demonstrates an improvement in classifying scenes with different levels of structural damage to buildings, roads, and infrastructures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。