研究视觉问答中模型与人类回答的差异,发现现有模型难以捕捉人类不确定性。
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
- 按人类回答分歧程度分层评估模型表现,引入新指标衡量与人类分布的匹配度
- 即使顶尖模型BEiT3也难以捕捉多标签回答分布,且常用校准方法反而加剧偏差
- 建议以人类不确定性为导向校准模型,适合关注模型可解释性与人机对齐的研究者
大型视觉语言模型在预测多人标注者的回答时经常表现不佳,尤其是在人类存在不确定性的场景下。本研究聚焦视觉问答(VQA)任务,全面评估当前最优模型与人类回答分布之间的相关性。我们根据人类回答分歧度(HUD)将样本划分为低、中、高三个层次,并不仅使用准确率,还引入三种新的与人类相关的评估指标,探究HUD对模型表现的影响。为更好对齐人类行为,我们验证了通用校准与人类校准的效果。结果表明,即便最先进的BEiT3模型仍难以捕捉多样人类回答中的多标签分布。此外,常用的以准确率为导向的校准方法会削弱BEiT3对HUD的建模能力,进一步拉大模型预测与人类分布的差距。相反,以人类分布为导向的校准能有效提升模型置信度与人类不确定性的对齐。研究揭示,当前对人类回答与模型预测一致性的问题仍被严重忽视,应成为未来研究的核心目标。
原文摘要 · Abstract (English)
Large vision-language models frequently struggle to accurately predict responses provided by multiple human annotators, particularly when those responses exhibit human uncertainty. In this study, we focus on the Visual Question Answering (VQA) task, and we comprehensively evaluate how well the state-of-the-art vision-language models correlate with the distribution of human responses. To do so, we categorize our samples based on their levels (low, medium, high) of human uncertainty in disagreement (HUD) and employ not only accuracy but also three new human-correlated metrics in VQA, to investigate the impact of HUD. To better align models with humans, we also verify the effect of common calibration and human calibration. Our results show that even BEiT3, currently the best model for this task, struggles to capture the multi-label distribution inherent in diverse human responses. Additionally, we observe that the commonly used accuracy-oriented calibration technique adversely affects BEiT3's ability to capture HUD, further widening the gap between model predictions and human distributions. In contrast, we show the benefits of calibrating models towards human distributions for VQA, better aligning model confidence with human uncertainty. Our findings highlight that for VQA, the consistent alignment between human responses and model predictions is understudied and should become the next crucial target of future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。