arXiv:2607.23921cs.CV2026-07中稿 · CGI

提出动态双层融合网络,精准定位图文问答中的答案区域。

DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

论文配图:DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
图 1 · 摘自论文原文
  • 用问题引导的动态区域模块融合图像与文本信息
  • 跨模态多尺度聚合提升像素与区域特征融合效果
  • 适合需要精细视觉定位的图文问答任务

视觉问答中的答案定位旨在根据自然语言问题,从图像中定位相关视觉区域,具有重要应用价值。本文提出动态双层视觉变压器融合网络(DDVT),设计问题引导的动态区域级模块(QGDR),通过ROI Align结合图像上下文与文本内容,实现对文本相关视觉内容的精确定位。同时引入跨模态多尺度聚合模块(CMA),增强像素级与区域级特征间的融合,有效定位与答案相关的视觉内容。最后将定位结果与文本特征融合,完成区域定位与答案生成。实验表明,该方法在多个主流基准上优于现有先进方法。

原文摘要 · Abstract (English)

Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.

视觉问答答案定位多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。