让图像搜索智能体像人一样记住上下文,提升多步推理能力
PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search

- 构建分层记忆系统:工作记忆、情景记忆和语义记忆协同推理
- 在DISBench上最高提升18.5%,有效减少上下文丢失和执行偏差
- 无需训练,适配多种多模态大模型,适合复杂图像搜索任务
深度图像搜索需要对时间、地点、事件关系等丰富上下文线索进行多步推理。然而,现有基于大语言模型的智能体大多无状态且被动响应,缺乏持久记忆以维持长时上下文或跨任务经验传递,常导致执行漂移与经验孤立。为此,我们提出PhotoCraft,一种无需训练的分层记忆系统,用于照片搜索智能体。受人类认知启发,PhotoCraft为多模态大模型赋予工作记忆、情景记忆和语义记忆,在推理过程中动态调用,以保持多步推理与答案生成中的逻辑一致性与知识可迁移性。在DISBench上的大量实验表明,PhotoCraft在多种多模态大模型骨干网络上均显著提升上下文感知检索性能,最高提升达18.5%,有效缓解了无记忆深度图像搜索的关键瓶颈,为构建可靠且通用的多模态搜索智能体提供了可行路径。
原文摘要 · Abstract (English)
Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and reactive, lacking persistent memory to maintain long-horizon context or transfer experience across tasks, which often leads to execution drift and experience isolation. To address these limitations, we propose PhotoCraft, a training-free, hierarchical memory system for photo-search agents. Inspired by human cognition, PhotoCraft equips MLLMs with working, episodic, and semantic memory, which are dynamically invoked during reasoning to preserve logical consistency and knowledge transferability throughout multi-step reasoning and answer generation. Extensive experiments on DISBench demonstrate that PhotoCraft consistently improves context-aware retrieval across diverse MLLM backbones, achieving gains of up to 18.5\% and effectively mitigating key bottlenecks in memoryless deep image search, offering a practical path toward reliable and generalizable multimodal search agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。