arXiv:2509.09721cs.CVcs.AI2025-09被引 5

用图文联合检索增强生成,提升灾后房屋损毁评估的准确率

A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval

  • 构建双分支编码结构,分别处理灾后图像与文本政策信息
  • 跨模态对齐提升检索精度,Top-1准确率提升9.6%
  • 适合保险理赔与应急资源规划场景使用

自然灾害后,准确评估住房损毁情况对于保险理赔响应和资源调配至关重要。本文提出一种新型多模态检索增强生成(MM-RAG)框架。在经典RAG架构基础上,设计双分支多模态编码结构:图像分支采用由ResNet与Transformer组成的视觉编码器,提取灾后建筑损毁特征;文本分支则利用BERT检索器对社交媒体帖子及保险条款进行向量化,并构建可检索的修复指数。为实现跨模态语义对齐,模型引入交叉注意力机制的跨模态交互模块。在生成模块中,通过模态注意力门控机制动态调控视觉证据与文本先验信息的作用。整个框架采用端到端训练,结合对比损失、检索损失与生成损失构成多任务优化目标,实现图像理解与政策匹配的协同学习。实验表明,该框架在检索准确率与损毁程度分类指标上均表现优异,其中Top-1检索准确率提升9.6%。

原文摘要 · Abstract (English)

After natural disasters, accurate evaluations of damage to housing are important for insurance claims response and planning of resources. In this work, we introduce a novel multimodal retrieval-augmented generation (MM-RAG) framework. On top of classical RAG architecture, we further the framework to devise a two-branch multimodal encoder structure that the image branch employs a visual encoder composed of ResNet and Transformer to extract the characteristic of building damage after disaster, and the text branch harnesses a BERT retriever for the text vectorization of posts as well as insurance policies and for the construction of a retrievable restoration index. To impose cross-modal semantic alignment, the model integrates a cross-modal interaction module to bridge the semantic representation between image and text via multi-head attention. Meanwhile, in the generation module, the introduced modal attention gating mechanism dynamically controls the role of visual evidence and text prior information during generation. The entire framework takes end-to-end training, and combines the comparison loss, the retrieval loss and the generation loss to form multi-task optimization objectives, and achieves image understanding and policy matching in collaborative learning. The results demonstrate superior performance in retrieval accuracy and classification index on damage severity, where the Top-1 retrieval accuracy has been improved by 9.6%.

多模态灾后评估RAG图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。