arXiv:2607.09822cs.CVcs.AI2026-07

用视觉记忆提升摄像头首位智能体的工具调用精准度。

Memory-Conditioned Tool Calling for Camera-First Visual Agents

  • 构建三层个人视觉记忆模型,动态调节工具选择与参数。
  • 移除记忆后工具相关性下降11.2%,整体效用降低9.7%。
  • 适合研究多工具协作与记忆增强型视觉智能体的开发者。

识别告诉智能体图像中有什么;而个人记忆则决定下一步该查询什么。在仅能发送图像的摄像头首位场景中,智能体必须自主生成查询。本文研究个人视觉记忆是否能改善智能体侧的工具选择与参数设定,从而实现更贴近用户的多工具查询。设计采用三层个人视觉记忆(档案、短期关注点、观测记录),每轮加载以引导基于大语言模型的工具调用循环,并引入冲突感知的写回机制,以更新后续捕捉的用户模型。在800张图像与合成记忆块组成的对照实验中,移除完整三层记忆使工具查询相关性绝对下降0.47分(5分制从4.21降至3.74,相对下降11.2%),端到端效用下降0.082(从0.842降至0.760,相对下降9.7%)。结果量化了在仅图像输入下,记忆对工具策略的调控作用,但未包含来自真实用户历史的多会话写回。

原文摘要 · Abstract (English)

Recognition tells an agent what is in an image; personal memory affects what is worth looking up next. In a camera-first setting the user can send only an image, so the agent must form the lookups. We study whether personal visual memory improves agent-side tool choice and tool arguments, and thereby more user-aligned multi-tool lookups. The design uses a three-layer personal visual memory (profile, short-term focus, observations) that is loaded on each turn to condition an LLM tool-calling loop under camera-first intake, and includes conflict-aware write-back intended to refresh the user model for later captures. On 800 images paired with synthetic memory blocks constructed for controlled ablation, removing the full three-layer memory block reduces tool-query relevance by 0.47 points absolute (4.21 -> 3.74 on a 5-point scale; 11.2% relative) and end-to-end utility by 0.082 absolute (0.842 -> 0.760; 9.7% relative). These results measure memory conditioning of tool policy under image-only intake with fixed synthetic blocks, not multi-session write-back from live user histories.

视觉记忆工具调用智能体摄像头首位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。