提出协同参数校准框架,提升视觉问答中检索与生成的交互效果。
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
- 设计统一框架,让检索器与生成器在训练和推理中共享参数知识。
- 在多个数据集上实现4.7%的准确率提升,基础大模型平均增益7.5%。
- 引入后期交互与反思机制,增强对问题与外部知识的细粒度理解。
基于知识的视觉问答(KB-VQA)系统通过从外部知识库检索知识来回答复杂的视觉接地问题。知识检索与答案生成均需对问题上下文和外部知识进行精确的多模态理解。然而,现有方法将两个阶段视为独立模块,训练中交互有限,阻碍了双向参数知识共享,最终导致性能不佳。为充分挖掘KB-VQA中的跨任务协同效应,我们提出一种统一的检索增强型VQA框架,实现协同参数知识校准。该框架能有效适配通用多模态预训练模型,完成细粒度、知识密集型任务,同时在训练与推理过程中使检索器与生成器协同增强并共享其参数知识。为提升对问题与外部文档的细粒度理解,我们在训练框架中引入后期交互机制。此外,还提出一种反思式回答机制,使模型能显式评估并优化其知识边界。所提方法在多项基准上表现优异,相比现有最优模型,回答准确率提升4.7%,基础多模态大模型(MLLMs)的VQA性能平均提升7.5%。
原文摘要 · Abstract (English)
Knowledge-based Vision Question Answering (KB-VQA) systems address complex visual-grounded questions with knowledge retrieved from external knowledge bases. The tasks of knowledge retrieval and answer generation tasks both necessitate precise multimodal understanding of question context and external knowledge. However, existing methods treat these two stages as separate modules with limited interaction during training, which hinders bi-directional parametric knowledge sharing, ultimately leading to suboptimal performance. To fully exploit the cross-task synergy in KB-VQA, we propose a unified retrieval-augmented VQA framework with collaborative parametric knowledge calibration. The proposed framework can effectively adapt general multimodal pre-trained models for fine-grained, knowledge-intensive tasks while enabling the retriever and generator to collaboratively enhance and share their parametric knowledge during both training and inference. To enhance fine-grained understanding of questions and external documents, we also integrate late interaction mechanism into the proposed training framework. Additionally, we introduce a reflective-answering mechanism that allows the model to explicitly evaluate and refine its knowledge boundary. Our approach achieves competitive performance against state-of-the-art models, delivering a significant 4.7\% improvement in answering accuracy, and brings an average 7.5\% boost in base MLLMs' VQA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。