轻量级多模态记忆系统,让手机眼镜记住你的日常
LightMem-Ego: Your AI Memory for Everyday Life

- 用分层记忆结构持续记录视觉音频流
- 支持对象查找、对话回忆等六类日常查询
- 可部署在手机眼镜,适合生活助理场景
移动与可穿戴设备上的个人AI助手通过视觉和音频流持续感知用户日常生活。但回答关于过往经历的问题,需要能持续积累、组织并检索长期经验的轻量级多模态记忆系统,这仍是挑战。为此,我们提出LightMem-Ego,一种面向日常生活的轻量级流式多模态记忆系统。该系统持续捕获第一人称视觉与音频流,在共享时间线上对齐,并组织为包含当前、短期和长期记忆的分层结构。给定用户查询时,LightMem-Ego动态路由至相应记忆层级,基于多模态证据生成答案。演示可在智能手机和AI眼镜上部署,支持对象查找、对话回忆、生活总结、习惯发现和个性化辅助。代码已开源:https://github.com/zjunlp/LightMem-Ego。
原文摘要 · Abstract (English)
Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously accumulate, organize, and retrieve long-term experiences, which remains challenging. To address this challenge, we present LightMem-Ego, a lightweight streaming multimodal memory system for everyday-life assistance. The system continuously captures egocentric visual and audio streams, aligns them on a shared timeline, and organizes them into a hierarchical memory consisting of current, short-term, and long-term memory. Given a user query, LightMem-Ego dynamically routes retrieval to the appropriate memory level and generates answers grounded in multimodal evidence. The demonstration can be deployed on smartphones and AI glasses, supporting object finding, conversation recall, life summarization, routine discovery, and personalized assistance. Code is available at https://github.com/zjunlp/LightMem-Ego.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。