从历史样本中推断图像构建的领域关系,提升VQA准确率
Abduction of Domain Relationships from Data for VQA
- 通过归纳推理从过往例子中推断图像元素间的关系
- 仅需少量示例即可显著提高问答准确率
- 适合缺乏领域数据的视觉问答场景
本文研究视觉问答(VQA)问题,其中图像和查询由缺乏领域数据的ASP程序表示。我们提出一种与现有知识增强技术正交且互补的方法,通过过往样本归纳出图像构建的领域关系。在建立归纳问题框架后,给出基线方法及实现,该方法显著提升了查询回答的准确性,且所需示例极少。
原文摘要 · Abstract (English)
In this paper, we study the problem of visual question answering (VQA) where the image and query are represented by ASP programs that lack domain data. We provide an approach that is orthogonal and complementary to existing knowledge augmentation techniques where we abduce domain relationships of image constructs from past examples. After framing the abduction problem, we provide a baseline approach, and an implementation that significantly improves the accuracy of query answering yet requires few examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。