评测智能体在个人文件中的多模态管理能力,揭示当前模型在用户画像和跨模态推理上的短板。
HippoCamp: Benchmarking Contextual Agents on Personal Computers
- 构建真实用户文件系统,评估智能体在个人环境中的上下文理解能力。
- 顶尖模型在用户画像任务上仅达48.3%准确率,长序列检索与跨模态推理表现差。
- 提供细粒度标注轨迹,支持对智能体失败原因逐步诊断,适合个人AI研究者使用。
我们提出HippoCamp,一个针对多模态文件管理的新型基准,用于评估智能体在以用户为中心环境中的能力。与聚焦网页交互、工具使用或通用场景自动化现有基准不同,HippoCamp模拟真实用户档案下的设备级文件系统,覆盖多样模态数据,包含超过2,000个真实文件,总数据量达42.4 GB。在此基础上,构建了581个问答对,用于评估智能体在搜索、证据感知和多步推理方面的能力。为支持细粒度分析,我们提供了46.1万条密集标注的结构化轨迹,实现逐步骤故障诊断。我们在HippoCamp上评估了多种先进多模态大语言模型(MLLMs)和代理方法。实验表明,即使最先进的商用模型在用户画像任务上也仅达到48.3%准确率,尤其在长时序检索和密集个人文件中的跨模态推理方面表现不佳。进一步的逐步诊断揭示,多模态感知与证据定位是主要瓶颈。HippoCamp揭示了当前智能体在真实用户环境中存在的关键局限,并为下一代个人AI助手的发展提供了坚实基础。
原文摘要 · Abstract (English)
We present HippoCamp, a new benchmark designed to evaluate agents' capabilities on multimodal file management. Unlike existing agent benchmarks that focus on tasks like web interaction, tool use, or software automation in generic settings, HippoCamp evaluates agents in user-centric environments to model individual user profiles and search massive personal files for context-aware reasoning. Our benchmark instantiates device-scale file systems over real-world profiles spanning diverse modalities, comprising 42.4 GB of data across over 2K real-world files. Building upon the raw files, we construct 581 QA pairs to assess agents' capabilities in search, evidence perception, and multi-step reasoning. To facilitate fine-grained analysis, we provide 46.1K densely annotated structured trajectories for step-wise failure diagnosis. We evaluate a wide range of state-of-the-art multimodal large language models (MLLMs) and agentic methods on HippoCamp. Our comprehensive experiments reveal a significant performance gap: even the most advanced commercial models achieve only 48.3% accuracy in user profiling, struggling particularly with long-horizon retrieval and cross-modal reasoning within dense personal file systems. Furthermore, our step-wise failure diagnosis identifies multimodal perception and evidence grounding as the primary bottlenecks. Ultimately, HippoCamp exposes the critical limitations of current agents in realistic, user-centric environments and provides a robust foundation for developing next-generation personal AI assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。