MGA-VQA通过图结构增强实现可解释的文档视觉问答
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
- 用图结构建模文本、空间与视觉特征的关联关系
- 在6个数据集上提升答案准确率与定位精度
- 适合需要可解释性的高安全文档分析场景
文档视觉问答(DocVQA)要求模型联合理解文本语义、空间布局和视觉特征。现有方法在显式建模空间关系、处理高分辨率文档时效率低、难以进行多跳推理,且可解释性差。我们提出MGA-VQA,一种多模态框架,融合了标记级编码、空间图推理、记忆增强推理和问题引导压缩。不同于以往黑箱模型,MGA-VQA引入可解释的基于图的决策路径和结构化记忆访问,提升推理透明度。在六个基准测试(FUNSD、CORD、SROIE、DocVQA、STE-VQA 和 RICO)上的评估表明,该模型在答案预测和空间定位方面均表现出更优的准确率与效率,且结果具有一致性提升。
原文摘要 · Abstract (English)
Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spatial relationship modeling, inefficiency with high-resolution documents, multi-hop reasoning, and limited interpretability. We propose MGA-VQA, a multi-modal framework that integrates token-level encoding, spatial graph reasoning, memory-augmented inference, and question-guided compression. Unlike prior black-box models, MGA-VQA introduces interpretable graph-based decision pathways and structured memory access for enhanced reasoning transparency. Evaluation across six benchmarks (FUNSD, CORD, SROIE, DocVQA, STE-VQA, and RICO) demonstrates superior accuracy and efficiency, with consistent improvements in both answer prediction and spatial localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。