arXiv:2511.12131cs.CVcs.AI2025-11AAAI

用物体属性描述提升大模型零样本视觉问答能力

OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description

  • 通过生成物体聚焦样本和全局描述,减少语言偏见影响
  • 在零样本场景下显著提升视觉问答准确率,超越现有方法
  • 适合需要鲁棒跨域推理的视觉理解任务研究者

大型语言模型(LLMs)在少样本或零样本视觉问答(VQA)中处理知识密集型问题方面已成关键工具。然而,其对大规模训练数据的依赖常导致语言偏见的继承,限制了模型表现:一方面,偏见使预测可靠性下降;另一方面,尽管具备强大推理能力,模型在分布外(OOD)场景下仍难以泛化。为此,我们提出物体属性描述增强器(OAD-Promoter),通过三个模块缓解语言偏见并提升领域迁移鲁棒性。其中,物体聚焦样本生成(OEG)模块生成全局描述与物体聚焦样本,结合全局与局部视觉线索增强输入信息并抑制偏见;记忆知识辅助(MKA)模块从存储示例中检索相关知识,支持未见领域的问答;OAD提示模块整合前序模块输出,优化大模型推理过程。实验表明,该方法在少样本/零样本设置下显著提升基于大模型的VQA性能,达到新的最先进水平。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become a crucial tool in Visual Question Answering (VQA) for handling knowledge-intensive questions in few-shot or zero-shot scenarios. However, their reliance on massive training datasets often causes them to inherit language biases during the acquisition of knowledge. This limitation imposes two key constraints on existing methods: (1) LLM predictions become less reliable due to bias exploitation, and (2) despite strong knowledge reasoning capabilities, LLMs still struggle with out-of-distribution (OOD) generalization. To address these issues, we propose Object Attribute Description Promoter (OAD-Promoter), a novel approach for enhancing LLM-based VQA by mitigating language bias and improving domain-shift robustness. OAD-Promoter comprises three components: the Object-concentrated Example Generation (OEG) module, the Memory Knowledge Assistance (MKA) module, and the OAD Prompt. The OEG module generates global captions and object-concentrated samples, jointly enhancing visual information input to the LLM and mitigating bias through complementary global and regional visual cues. The MKA module assists the LLM in handling OOD samples by retrieving relevant knowledge from stored examples to support questions from unseen domains. Finally, the OAD Prompt integrates the outputs of the preceding modules to optimize LLM inference. Experiments demonstrate that OAD-Promoter significantly improves the performance of LLM-based VQA methods in few-shot or zero-shot settings, achieving new state-of-the-art results.

视觉问答大模型零样本偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。