用知识图谱增强灾后图像描述,让生成内容更准确专业。
VLCE: A Knowledge-Enhanced Framework for Image Description in Disaster Assessment
- 结合知识图谱与视觉模型,用1566个灾损专有词优化描述
- 在救援网数据集上,95%的评价偏好知识增强版本
- 适合灾情评估、应急响应等需要精准描述的场景
通用视觉语言模型(如LLaVA和QwenVL)生成的灾后影像描述缺乏领域术语和可操作细节。我们提出视觉语言描述增强框架VLCE,将ConceptNet和WordNet的外部语义知识融入灾后卫星与无人机影像的描述生成过程。VLCE分两阶段:第一阶段由基线VLM根据YOLOv8检测结果生成初始描述;第二阶段通过CNN-LSTM或层次化跨模态Transformer,利用包含1,566个领域相关词汇的增强词表进行精修。我们在两个灾后基准测试上评估:xBD(卫星,6,369张图,3类损伤)和RescueNet(无人机,4,494张图,12类损伤),使用CLIPScore衡量语义对齐度,InfoMetIC衡量信息量。在RescueNet上使用Transformer解码器时,知识图谱增强的VLCE在InfoMetIC上优于QwenVL基线95.33%,在CLIPScore上为73.64%。定性分析显示,无知识图谱时描述存在幻觉、重复和语义混乱;加入知识后,描述保持事实一致性与领域适配性。
原文摘要 · Abstract (English)
General-purpose vision-language models (VLMs) such as LLaVA and QwenVL produce descriptions of disaster imagery that lack domain-specific vocabulary and actionable detail. We propose the Vision-Language Caption Enhancer (VLCE), a framework that integrates external semantic knowledge from ConceptNet and WordNet into the caption generation process for post-disaster satellite and UAV imagery. VLCE operates in two stages: first, a baseline VLM generates an initial caption conditioned on YOLOv8 object detections; second, a knowledge-enriched sequential model, a CNN-LSTM or a hierarchical cross-modal Transformer, refines the caption using a vocabulary augmented with 1,566 domain-relevant terms extracted from knowledge graphs. We evaluate VLCE on two disaster benchmarks: xBD (satellite, 6,369 images, 3 damage classes) and RescueNet (UAV, 4,494 images, 12 damage classes), using CLIPScore for semantic alignment and InfoMetIC for informativeness. On RescueNet with the Transformer decoder, VLCE with knowledge graph enrichment produces captions preferred over QwenVL baselines in 95.33% of image pairs on InfoMetIC and 73.64% on CLIPScore. Qualitative analysis shows that without knowledge graph integration, generated captions exhibit hallucinations, word repetition, and semantic incoherence, whereas knowledge-enriched captions maintain factual consistency and domain-appropriate vocabulary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。