提出可编辑状态的视觉实例绑定,解决多轮个性化定位中证据利用难题。
SCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization

- 构建状态驱动的多轮定位框架,通过事件序列追踪目标状态变化。
- 引入过渡树与视觉证据自适应机制,使定位准确率提升至70.27%。
- 特别擅长处理非最新或回滚状态的定位,适合研究多轮交互系统。
我们提出可编辑状态条件下的视觉实例绑定(SCVIB),一种多轮定位设置,其中多个支持实例在不同轮次中引入,协议定义的状态事件决定最终目标。该设置包含1,050对人工验证的支持-查询样本和1,500个覆盖五个视觉领域、三种难度等级和四类目标状态依赖关系的剧集。直接无序推理仅达60.13% [email protected],表明确定最终参考并不等于有效利用对应视觉证据进行查询端定位。为此,我们提出TT-VG(过渡树视觉定位),结合目标状态转换树(TSTT)与视觉证据定位适应(VEGA)。TSTT将可见交互编译为协议定义事件,在版本化目标状态下执行,并将最终查询引用解析至对应支持证据。基于轨迹衍生的同实例对进行适配,VEGA使用视觉证据包实现支持条件下的实例定位。TT-VG达到70.27% [email protected];在匹配目标解析条件下,优于最强对比方法16.20点。相较于直接推理,性能提升在反近期性与回滚任务中最大,这两类任务需路由至非最新或已恢复的支持证据。结果共同确立了SCVIB作为可控测试平台,并凸显在多轮个性化定位中有效利用解析出的支持证据是核心挑战。
原文摘要 · Abstract (English)
We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the final target. We instantiate this setting as SCVIB, comprising 1,050 manually verified support--query base pairs and 1,500 episodes spanning five visual domains, three difficulty levels, and four target-state dependency groups. Direct Seq-free inference reaches only 60.13\% [email protected], indicating that resolving the final reference does not ensure effective use of the corresponding visual evidence for query-side localization. We address this gap with TT-VG (Transition-Tree Visual Grounding), which combines a Target-State Transition Tree (TSTT) with Visual Evidence Grounding Adaptation (VEGA). TSTT compiles the visible interaction into protocol-defined events, executes them over versioned target states, and resolves the final-query reference to the corresponding support evidence. Adapted on trajectory-derived same-instance pairs, VEGA performs support-conditioned grounding of the resolved instance using a Visual Evidence Package. TT-VG reaches 70.27\% [email protected]; under matched target resolution, VEGA exceeds the strongest comparison method by 16.20 points. Gains over direct inference are largest on Counter-Recency and Rollback, which require routing to non-latest or restored support evidence. Together, these results establish SCVIB as a controlled testbed and highlight the effective use of resolved support evidence for query-side same-instance localization as a central challenge in multi-turn personalized localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。