通过自我反思与跨模型验证提升视觉问答可靠性
Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification
- 双路径架构:融合模型特征与问答嵌入评估可信度,外接参考模型交叉验证事实
- 在ICCV-CLVL 2025可靠问答挑战中获Φ₁₀₀ 39.64、100-AUC 97.22,排名第一
- 适合关注大模型幻觉问题与答案可信度的开发者和研究者
视觉语言模型(VLMs)在视觉问答(VQA)任务中展现出巨大潜力,但其易受幻觉影响,导致自信却错误的回答,严重损害答案可靠性。为此,我们提出双评估框架DAVR,融合自我反思与跨模型验证,实现全面的不确定性估计。DAVR采用双路径结构:一条路径利用双选择器模块,通过融合VLM隐层特征与问答嵌入评估回答可信度;另一条路径部署外部参考模型进行事实交叉核查,以缓解幻觉问题。在ICCV-CLVL 2025可靠问答挑战中,DAVR取得Φ₁₀₀ 39.64、100-AUC 97.22的领先成绩,位居第一,验证了其在提升VLM回答可信性方面的有效性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated significant potential in Visual Question Answering (VQA). However, the susceptibility of VLMs to hallucinations can lead to overconfident yet incorrect answers, severely undermining answer reliability. To address this, we propose Dual-Assessment for VLM Reliability (DAVR), a novel framework that integrates Self-Reflection and Cross-Model Verification for comprehensive uncertainty estimation. The DAVR framework features a dual-pathway architecture: one pathway leverages dual selector modules to assess response reliability by fusing VLM latent features with QA embeddings, while the other deploys external reference models for factual cross-checking to mitigate hallucinations. Evaluated in the Reliable VQA Challenge at ICCV-CLVL 2025, DAVR achieves a leading $Φ_{100}$ score of 39.64 and a 100-AUC of 97.22, securing first place and demonstrating its effectiveness in enhancing the trustworthiness of VLM responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。