arXiv:2608.16805cs.CVcs.AI2026-08

发现大模型常把属性错配给同类别物体,提出可量化该错误的新基准。

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

论文配图:Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
图 1 · 摘自论文原文
  • 构建新基准InstaBind-Lite,精准测量属性错配问题。
  • 开源模型平均错配率19.84%,商用接口仅7.55%,但都被整体准确率掩盖。
  • 80%以上错配来自相邻物体,定位和实例优先干预效果有限。

大型视觉语言模型虽能识别密集场景中的物体与属性,却常将属性错误分配给同类别对象。通用VQA准确率将此类回答标记为错误,而物体幻觉指标可能误判对象与属性均为图像支持;二者均无法揭示属性转移。本研究将此盲点定义为密集同类别属性错配(DSCAM),提出InstaBind-Lite——一个可控的基准测试,可直接测量该问题。其包含524张图像、529组3-6个同类别实体、1773个框定实例、有序邻近关系、可区分的颜色类属性及四类互补问题,生成9580个可确定评估的问题。与现有协议不同,源实例标注将无依据生成与识别失败,与从可见实体复制属性区分开。绑定特定度量进一步量化转移频率、邻近性、序数距离及干预影响。在五个开源和两个商用/API模型中,开源系统平均错配率为19.84%,商用系统为7.55%;这些错误被整体准确率隐藏。在可识别转移中,分别有80.70%和81.51%源自相邻实例。定位与实例优先干预对部分模型有效,但非普适解法。InstaBind-Lite因此将原本无法区分的错误答案转化为可溯源的故障类别,并检验传统基准无法衡量的可靠性维度:模型不仅知道可见物,更清楚每个属性属于哪个实例。

原文摘要 · Abstract (English)

Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents InstaBind-Lite, a controlled benchmark that makes it directly measurable. Its 524 images contain 529 curated groups of 3-6 same-class entities, 1773 boxed instances, ordered neighbors, distinguishable color-like attributes, and four complementary question levels, yielding 9580 deterministically evaluated questions. Unlike existing protocols, source-instance annotations separate unsupported generation and recognition failure from an attribute copied from another visible entity. Binding-specific metrics further quantify transfer frequency, adjacency, ordinal distance, and intervention effects. Across five open-source and two commercial/API models, the open-source systems average 19.84% Misbinding Rate and the API systems 7.55%; these errors are hidden by aggregate accuracy. Among identifiable transfers, 80.70% and 81.51%, respectively, originate from adjacent instances. Localization and instance-first interventions help selected models but are not universal remedies. InstaBind-Lite therefore turns previously undifferentiated wrong answers into source-identifiable failure categories and tests a reliability dimension that conventional benchmarks cannot determine: whether a model knows not only what is visible, but which instance owns each attribute.

视觉语言模型属性错配模型评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。