为个人相册设计能理解长期视觉记忆的对话式智能助手
Personal AI Agent for Camera Roll VQA

- 构建分层记忆结构,高效检索用户多年积累的海量照片
- 在50名用户3.1万张图片上实现2500个问答的准确回答
- 适合做个性化视觉记忆、长期上下文理解的研究者
我们研究了个人相册视觉问答任务。在此场景中,对话式AI助手可访问用户个人相册,根据问题检索相关照片,回答从简单事实(如“我昨天试吃的食物名字?”)到开放性问题(如“推荐一些我没吃过的菜”)。由于个人相册内容庞大(覆盖数年、数百至数千张照片),成功的助手需理解长期、高度个性化的视觉内容流,以精准定位信息。为此,我们收集并人工标注了模拟真实使用场景的问题,最终形成camroll数据集:包含50名用户、31,476张图像和2,500组问答对。我们进一步设计camroll-agent,一个配备分层记忆和最小工具集的对话式智能体,实现对大型个性化视觉记忆的高效导航。实验表明,camroll-agent优于多个基线及长上下文理解型智能体系统。结果揭示:个人视觉记忆中的持续性、视觉细节与用户特定上下文,要求与标准长文本记忆不同的方法。
原文摘要 · Abstract (English)
We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera roll and retrieve relevant photos to answer queries, ranging from simple factual questions (e.g., ``Name of the food I tried yesterday?'') to more open-ended ones (e.g., ``Recommend some dishes I have never eaten before''). Given the vast nature of the personal camera roll (i.e., multiple years, hundreds to thousands of photos), a successful AI assistant needs to understand a long-horizon, highly personalized visual content stream in order to navigate and locate the correct and/or relevant information. To support this, we collect and manually annotate questions that mimic real-world usage. The final dataset, camroll, contains 50 users, 31,476 images, and 2,500 QA pairs. We further design camroll-agent, a conversational AI agent equipped with hierarchical memory and a minimal set of tools for efficient navigation over large, personalized visual memory. Experimental results show that camroll-agent outperforms numerous baselines and methods for long-context understanding AI agents system. Together, the camroll dataset and camroll-agent highlight the gap in AI agents' long-context reasoning: personalized visual memory requires different approaches from standard long-context textual memory, especially when consistency, visual details, and user-specific context are present.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。