arXiv:2605.12497cs.CV2026-05被引 2

让AI通过网络搜索理解图像,精准定位隐藏目标。

From Web to Pixels: Bringing Agentic Search into Visual Perception

论文配图:From Web to Pixels: Bringing Agentic Search into Visual Perception
图 1 · 摘自论文原文
  • 设计搜图流程,先用网络找目标信息再匹配像素
  • 在120张图上实现最高开源精度,错误多因信息获取失败
  • 适合研究视觉与外部知识融合的学者

视觉感知将高层语义理解与像素级感知连接,但现有方法多假设关键证据已存在于图像或模型知识中。本文研究更贴近现实的开放世界场景:需先从外部事实、近期事件、长尾实体或多跳关系中解析出可见物体的身份,才能定位。为此提出「感知深度搜索」范式,并构建了以对象为中心的WebEye基准,包含可验证证据、知识密集型问题、精确框/掩码标注及三种任务视图:基于搜索的定位、分割和VQA。WebEye含120张图像、473个标注实例、645个唯一问答对和1927个任务样本。进一步提出Pixel-Searcher,一种从搜索到像素的智能体工作流,用于解析隐藏目标身份并绑定至框、掩码或答案。实验表明,Pixel-Searcher在三项任务中均达到最强开源性能,失败主要源于证据获取、身份解析和视觉实例绑定环节。

原文摘要 · Abstract (English)

Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for identifying a target is already in the image or frozen model knowledge. We study a more practical yet harder open-world case where a visible object must first be resolved from external facts, recent events, long-tail entities, or multi-hop relations before it can be localized. We formalize this challenge as Perception Deep Research and introduce WebEye, an object-anchored benchmark with verifiable evidence, knowledge-intensive queries, precise box/mask annotations, and three task views: Search-based Grounding, Search-based Segmentation, and Search-based VQA. WebEyes contains 120 images, 473 annotated object instances, 645 unique QA pairs, and 1,927 task samples. We further propose Pixel-Searcher, an agentic search-to-pixel workflow that resolves hidden target identities and binds them to boxes, masks, or grounded answers. Experiments show that Pixel-Searcher achieves the strongest open-source performance across all three task views, while failures mainly arise from evidence acquisition, identity resolution, and visual instance binding.

视觉感知网络搜索智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。