让AI记住用户照片里的私人信息,比纯文字记忆更准。
Personal Visual Memory from Explicit and Implicit Evidence

- 用对话上下文解析图片中的身份与归属关系
- 在新基准上显著超越现有记忆系统
- 适合需要懂用户私密信息的个性化AI
长期记忆对个性化AI代理越来越重要,但现有评测和方法仍以文本为中心。即使包含图像,用户相关信息通常仅靠文本即可恢复,多数记忆系统将图像简化为通用描述。然而,图像常携带文本难以表达的个人化信息:显性证据如反复出现的用户关联实体,隐性证据如通过视觉或跨模态线索推断的潜在事实。我们提出一个针对两类证据的个人视觉记忆基准,并设计了VisualMem——一种混合视觉-文本架构,通过结构化个人视觉记忆模块增强文本记忆后端。VisualMem不将图像压缩为摘要,而是利用对话上下文识别身份、所有权及持久用户事实。实验表明,VisualMem在新基准上显著优于现有记忆系统,同时在标准文本记忆基准上保持竞争力,证明个人视觉记忆是个性化AI长期记忆中独特且关键的组成部分。
原文摘要 · Abstract (English)
Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when images are included, the user-specific information needed for later questions is typically recoverable from text alone, and most memory systems reduce image turns to generic captions. Yet images often carry personal information that text rarely states -- both explicit evidence, such as recurring user-associated entities, and implicit evidence, such as latent user facts inferred from visual or multimodal cues. We introduce a benchmark for personal visual memory that targets both forms of evidence, and propose VisualMem, a hybrid visual--text architecture that augments a text-memory backend with a structured personal visual memory module. Rather than collapsing images into captions, VisualMem uses conversational context to resolve identity, ownership, and durable user facts. Experiments show that VisualMem substantially outperforms prior memory systems on our benchmark while remaining competitive on standard text-memory benchmarks, indicating that personal visual memory is a distinct and important component of long-term memory for personalized AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。