用概率方法重加权数据,提升视觉问答的检索精度。
Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering

- 将文档重要性建模为潜在变量,自适应调整负样本权重
- 在7个VQA基准上,3种检索器均实现精度提升
- 适合需要精准知识检索的多模态问答研究者
基于知识的视觉问答依赖多模态检索器从外部证据中获取信息。然而,现有对比学习方法通常将所有不匹配的查询-文档对视为同等信息量的负样本,这存在问题:许多未匹配文档仍可能语义相关或部分有用。本文提出贝叶斯数据重加权,一种概率框架,将查询-文档重要性建模为潜在变量,通过后验推断自适应降低可能的假负样本权重。在共轭先验下采用闭式后验更新与随机EM优化,该方法在三种检索器和七个知识型VQA基准上均持续提升检索准确率。
原文摘要 · Abstract (English)
Multimodal retrievers are essential for knowledge-based visual question answering, where they retrieve external evidence for image-question pairs. However, existing contrastive training methods typically treat all unmatched query-document pairs as equally informative negatives, which is problematic because many unmatched documents may still be semantically relevant or partially useful. We propose Bayesian Data Reweighting, a probabilistic framework that models query-document importance as latent variables and adaptively infers posterior weights to downweight likely false negatives. With closed-form posterior updates under conjugate priors and stochastic EM optimization, our method consistently improves retrieval accuracy across three retrievers and seven knowledge-based VQA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。