arXiv:2606.20736cs.CV2026-06

让视觉问答评测永远新鲜:动态重置图像关键细节,防止模型靠记忆得分。

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

论文配图:REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation
图 1 · 摘自论文原文
  • 评测时动态重置图像中的答案线索,确保模型靠理解而非记忆答题。
  • 在V*Bench上,前沿模型原题得分比重置后高9.5至18.8个百分点。
  • 适合关注真实视觉理解能力评估的研究者与评测团队使用。

静态视觉问答(VQA)基准测试会迅速过时:一旦数据泄露至训练语料,分数可能反映的是记忆而非真实视觉能力,从而掩盖实际进展。重建高质量基准如V*Bench需大量人工标注,但每次发布后又很快成为泄露样本。我们提出ReKey,一种实时基准协议,在评估时随机重生成图像中承载答案的局部细节(即视觉键)。通过人工验证的编辑区域,ReKey可生成带有新答案、基于结构的标签以及可控视觉搜索难度的新实例。在V*Bench上,重生成的基准揭示了八种前沿视觉语言模型(VLMs)的分数显著下降:原题得分比重置版本高出9.5至18.8个百分点。通过使视觉键可再生,ReKey能随模型和训练数据演进而持续保持评测新鲜度。

原文摘要 · Abstract (English)

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus obscuring real progress. Rebuilding high-quality benchmarks such as V*Bench requires substantial human annotation, yet each static release can quickly become another leaked artifact. We propose ReKey, a live benchmark protocol that randomly regenerates the answer-bearing local detail, or visual key, in real images at evaluation time. Using human-validated edit slots, ReKey samples fresh instances with new answers, construction-grounded labels, and controlled visual-search difficulty. On V*Bench, the ReKey regenerated benchmark reveals a sharp score jump across eight frontier vision-language models (VLMs): The original items score 9.5--18.8 percentage points higher than the regenerated variants. By making the visual key renewable, ReKey keeps evaluation fresh as models and training data evolve.

视觉问答评测基准抗记忆动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。