arXiv:2608.08663cs.HCcs.CV2026-08

让机器像人一样通过对话逐步建立指代共识,实现精准视觉定位。

A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

论文配图:A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data
图 1 · 摘自论文原文
  • 用三个可追踪的绑定集合显式记录对话中的指代表达状态
  • 单次指令下定位正确目标的准确率达83.56%,接近人类水平
  • 框架透明可审计,适合研究语言与视觉协同机制的学者

人类通过反复互动对新奇难描述物体达成共享命名,这一过程称为词汇同调(lexical entrainment)。当前主流视觉-语言模型在该任务上表现不佳:无法缩短指称、复用有效表达,也无法维持对话中的稳定共识状态。本文提出一种动态语义框架,将共识状态外化为三个显式可查的参照物绑定集(Γ, Ξ, Ω),并由动态语义上下文变更规则更新。符号层基于轻量级感知对齐流水线,通过SIFT同形变换和通用质量指数(UQI)将噪声人类指称映射至众包图像。在斯坦福重复指称游戏数据集(15,000+条导演-匹配者语句,使用抽象七巧板刺激物)上,框架在单次导演话语下将正确目标置于前5候选中的比例达83.56%。人类匹配者在相同数据集上的准确率为77%-80%。此外,在排除明显邻近图像的保留条件下评估,提供更保守的接地信号测量。消融实验分离了SIFT对齐、UQI、查询预处理和图像增强的贡献。核心贡献在于:透明可审计的符号层逐轮恢复词汇同调结构,搭配可逐项分析行为的感知通道。同时详细讨论了框架不具有的能力:非交互式、未与导演闭环反馈、其检索驱动的感知通道存在可量化边界泄漏效应。

原文摘要 · Abstract (English)

Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents that they do not shorten references, reuse successful expressions, or maintain stable pact state across turns. We present a framework that addresses the gap by externalizing pact state into three explicit, inspectable sets of referent-object bindings ($Γ, Ξ, Ω$), updated by a dynamic-semantics context-change rule. The symbolic layer sits on top of a lightweight perceptual-alignment pipeline that grounds noisy human referring expressions in crowd-sourced imagery via SIFT homographies and the Universal Quality Index. Evaluated on the Stanford Repeated Reference Game corpus (over 15{,}000 director-matcher utterances on abstract tangram stimuli), the framework places the correct target in its top-5 hypothesis set 83.56% of the time from a single director utterance. Human matcher top-1 accuracy on the same corpus is approximately 77-80%. We also report results on a held-out condition in which obvious tangram-adjacent images are excluded from the retrieved set, which provides a more conservative measurement of the grounding signal. Ablations isolate the contribution of each component: SIFT alignment, UQI, query preprocessing, and image augmentation. The central contribution is the combination: a transparent, auditable symbolic layer that recovers the structure of lexical entrainment turn by turn, paired with a perceptual channel whose behavior can be examined ablation by ablation. We also discuss in detail what the framework does not do. It is not interactive, it does not close the loop with the director, and its retrieval-driven perceptual channel is vulnerable to a class of leakage effects that we quantify and bound rather than wave away.

视觉定位语言理解对话系统可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。