arXiv:2608.10692cs.CLcs.AI2026-08

评测大模型在分散个人数据中完成指令的能力,发现其定位信息错误率高达79%。

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

论文配图:SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
图 1 · 摘自论文原文
  • 构建覆盖10个应用的250项任务,测试大模型跨应用推理与整合能力。
  • 最佳模型准确率仅57.3%,最差模型低至16.4%,信息定位失误占失败的79%。
  • 适合研究移动助手、多源信息融合与大模型可信推理的学者与开发者。

大语言模型(LLMs)正越来越多地被用作移动助手,但其在利用分散于多个应用程序中的个人信息来完成用户指令方面仍面临挑战,主要原因在于缺乏专门的评估基准。为此,我们提出了SPIEval,一个基于五种认知能力(即推理、歧义消解、信息整合、偏好推断和多意图分解)的人工标注基准。SPIEval包含250项任务,覆盖4,335条个人记录,分布在10个应用中,并支持通过21种工具进行多轮交互。分析表明,该基准涵盖多样化场景、复杂任务、信息分散、可控环境及可验证结果。我们评估了九个代表性大模型,发现性能提升空间巨大:表现最好的模型GPT-5.5 (xhigh) 准确率为57.3%,最差模型仅为16.4%。进一步分析显示,79%的失败源于信息定位不准,模型常选择看似合理但错误的信息,而非继续检索验证。此外,少于2%的检索动作使用高级搜索方法,且各模型间搜索效率差异显著。这些发现揭示了当前基于大模型的移动助手存在根本性局限,推动该方向未来研究。数据与代码已公开于https://huggingface.co/datasets/Junjie-Ye/SPIEval。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

大模型评测移动助手信息整合多源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。