让电脑像人一样凭记忆找文件,支持时间地点活动多维查询。
Indaleko: The Unified Personal Index
- 构建统一索引图谱,融合时间空间活动元数据
- 3100万文件160TB数据下查询响应低于1秒
- 适合需要按记忆线索找资料的用户
个人信息检索在系统忽略人类记忆机制时失效。现有平台强制在孤立数据孤岛中进行关键词搜索,而人类自然通过何时、何地、何种情境等情景线索回忆信息。本文提出统一个人索引(UPI)架构,弥合这一根本差距。Indaleko原型在覆盖8个存储平台、总计3100万文件、160TB数据的测试集上验证了可行性。通过将时间、空间和活动元数据整合至统一图数据库,Indaleko支持‘去年春天在会议场地附近的照片’等自然语言查询,现有系统无法处理此类多维问题。实现亚秒级查询响应,消除跨平台搜索碎片化,并在明确记忆模式下保持完美精度。与谷歌云盘、OneDrive、Dropbox、Windows搜索对比评估显示,所有商业系统在基于记忆的查询中失败,返回海量无上下文筛选的结果。而Indaleko成功处理结合时间、位置与活动模式的复合查询。可扩展架构支持新数据源快速接入(每提供商10分钟至10小时),并通过基于UUID的语义解耦保障隐私。UPI架构融合认知理论与分布式系统设计,通过Indaleko原型与严格评估得以验证。该工作将个人信息检索从关键词匹配转变为记忆对齐发现,为现有数据提供即时价值,并为未来情境感知系统奠定基础。
原文摘要 · Abstract (English)
Personal information retrieval fails when systems ignore how human memory works. While existing platforms force keyword searches across isolated silos, humans naturally recall through episodic cues like when, where, and in what context information was encountered. This dissertation presents the Unified Personal Index (UPI), a memory-aligned architecture that bridges this fundamental gap. The Indaleko prototype demonstrates the UPI's feasibility on a 31-million file dataset spanning 160TB across eight storage platforms. By integrating temporal, spatial, and activity metadata into a unified graph database, Indaleko enables natural language queries like "photos near the conference venue last spring" that existing systems cannot process. The implementation achieves sub-second query responses through memory anchor indexing, eliminates cross-platform search fragmentation, and maintains perfect precision for well-specified memory patterns. Evaluation against commercial systems (Google Drive, OneDrive, Dropbox, Windows Search) reveals that all fail on memory-based queries, returning overwhelming result sets without contextual filtering. In contrast, Indaleko successfully processes multi-dimensional queries combining time, location, and activity patterns. The extensible architecture supports rapid integration of new data sources (10 minutes to 10 hours per provider) while preserving privacy through UUID-based semantic decoupling. The UPI's architectural synthesis bridges cognitive theory with distributed systems design, as demonstrated through the Indaleko prototype and rigorous evaluation. This work transforms personal information retrieval from keyword matching to memory-aligned finding, providing immediate benefits for existing data while establishing foundations for future context-aware systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。