让模型像人一样在视觉流中找图,靠上下文推理而非单张图片匹配。
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
- 将图像检索变为多步推理的自主探索任务,利用视觉历史中的上下文线索。
- 构建了包含复杂时序依赖的DISBench基准,现有模型在其中表现显著下降。
- 提出人机协同方法生成带上下文的查询,提升数据构建效率与质量。
现有多模态检索系统擅长语义匹配,但默认查询与图像的相关性可独立衡量。这一范式忽略了真实视觉流中信息分布于时间序列中的丰富依赖关系。为此,我们提出DeepImageSearch,一种新型代理范式,将图像检索重构为自主探索任务:模型需在原始视觉历史中规划并执行多步推理,基于隐含上下文线索定位目标。我们构建了DISBench——一个基于互联视觉数据的挑战性基准。为解决生成依赖上下文查询的可扩展性难题,提出人机协同流水线,利用视觉-语言模型挖掘潜在时空关联,再经人工验证,有效降低上下文发现成本。此外,设计了一个模块化代理框架,配备细粒度工具和双记忆系统,支持长时程导航。大量实验表明,DISBench对当前最先进模型构成严峻挑战,凸显下一代检索系统引入代理式推理的必要性。
原文摘要 · Abstract (English)
Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This paradigm overlooks the rich dependencies inherent in realistic visual streams, where information is distributed across temporal sequences rather than confined to single snapshots. To bridge this gap, we introduce DeepImageSearch, a novel agentic paradigm that reformulates image retrieval as an autonomous exploration task. Models must plan and perform multi-step reasoning over raw visual histories to locate targets based on implicit contextual cues. We construct DISBench, a challenging benchmark built on interconnected visual data. To address the scalability challenge of creating context-dependent queries, we propose a human-model collaborative pipeline that employs vision-language models to mine latent spatiotemporal associations, effectively offloading intensive context discovery before human verification. Furthermore, we build a robust baseline using a modular agent framework equipped with fine-grained tools and a dual-memory system for long-horizon navigation. Extensive experiments demonstrate that DISBench poses significant challenges to state-of-the-art models, highlighting the necessity of incorporating agentic reasoning into next-generation retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。