arXiv:2601.16394cs.CVcs.AI2026-01

用信息熵找关键点,让视觉推理更准

ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation

  • 基于信息熵从粗框中发现高信息量候选点
  • 在四个数据集上达到最新最好效果
  • 适合需要少样本精准分割的场景

指代表达分割(RES)是一项核心的视觉-语言分割任务,通过自然语言表达实现目标的像素级理解,支持人机交互和增强现实等应用。尽管基于多模态大模型(MLLM)的方法取得进展,现有方法仍存在两大缺陷:一是MLLM生成的粗略边界框导致冗余或非区分性点提示;二是普遍依赖文本坐标推理,难以区分视觉相似干扰项。为此,本文提出新型框架ResAgent,融合熵基点发现(EBD)与视觉推理(VBR)。EBD通过建模边界框内的空间不确定性,将点选择视为信息最大化过程,识别高信息候选点;VBR通过联合视觉-语义对齐验证点的正确性,摒弃纯文本坐标推断以提升鲁棒性。该框架采用从粗到精流程:边界框初始化、熵引导点发现、视觉验证与掩码解码。在RefCOCO、RefCOCO+、RefCOCOg和ReasonSeg四个基准数据集上的实验表明,ResAgent在所有数据集上均达到新最佳性能,证明其能以最少提示生成准确且语义一致的分割掩码。

原文摘要 · Abstract (English)

Referring Expression Segmentation (RES) is a core vision-language segmentation task that enables pixel-level understanding of targets via free-form linguistic expressions, supporting critical applications such as human-robot interaction and augmented reality. Despite the progress of Multimodal Large Language Model (MLLM)-based approaches, existing RES methods still suffer from two key limitations: first, the coarse bounding boxes from MLLMs lead to redundant or non-discriminative point prompts; second, the prevalent reliance on textual coordinate reasoning is unreliable, as it fails to distinguish targets from visually similar distractors. To address these issues, we propose \textbf{\model}, a novel RES framework integrating \textbf{E}ntropy-\textbf{B}ased Point \textbf{D}iscovery (\textbf{EBD}) and \textbf{V}ision-\textbf{B}ased \textbf{R}easoning (\textbf{VBR}). Specifically, EBD identifies high-information candidate points by modeling spatial uncertainty within coarse bounding boxes, treating point selection as an information maximization process. VBR verifies point correctness through joint visual-semantic alignment, abandoning text-only coordinate inference for more robust validation. Built on these components, \model implements a coarse-to-fine workflow: bounding box initialization, entropy-guided point discovery, vision-based validation, and mask decoding. Extensive evaluations on four benchmark datasets (RefCOCO, RefCOCO+, RefCOCOg, and ReasonSeg) demonstrate that \model achieves new state-of-the-art performance across all four benchmarks, highlighting its effectiveness in generating accurate and semantically grounded segmentation masks with minimal prompts.

指代分割视觉推理点提示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。