arXiv:2409.08250cs.HCcs.AI2024-09中稿 · the 2025 CHI Confe…被引 20

让手机记忆会说话:跨多段记忆推理,回答复杂个人问题

OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering

  • 用多段记忆整合上下文,增强单条记录的语义
  • 问答准确率达71.5%,比传统检索系统胜出或平局74.5%时间
  • 适合需要理解生活片段关联的人,如日记管理、回忆辅助

人们常通过照片、截图和视频记录生活。现有AI工具虽能用自然语言查询这些数据,但仅支持检索单一信息(如图片中物体),难以处理涉及多个记忆关联的复杂问题(如事件顺序)。我们开展为期一个月的日记研究,收集真实用户提问,并构建了整合记忆所需上下文信息的分类体系。随后提出OmniQuery系统,可回答需提取与推断上下文的个人记忆类问题。该系统通过整合多条关联记忆中的分散上下文,增强原始记录;给定问题后,检索增强后的记忆并利用大语言模型生成带引用的答案。在人工评估中,其准确率达到71.5%,在74.5%的测试中表现优于或持平于传统RAG系统。

原文摘要 · Abstract (English)

People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like certain objects in photos, and struggle with answering more complex queries that involve interpreting interconnected memories like sequential events. We conducted a one-month diary study to collect realistic user queries and generated a taxonomy of necessary contextual information for integrating with captured memories. We then introduce OmniQuery, a novel system that is able to answer complex personal memory-related questions that require extracting and inferring contextual information. OmniQuery augments individual captured memories through integrating scattered contextual information from multiple interconnected memories. Given a question, OmniQuery retrieves relevant augmented memories and uses a large language model (LLM) to generate answers with references. In human evaluations, we show the effectiveness of OmniQuery with an accuracy of 71.5%, outperforming a conventional RAG system by winning or tying for 74.5% of the time.

记忆增强多模态个性化问答上下文推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。