首个模拟真实手机使用场景的智能助手评测基准
iOSWorld: A Benchmark for Personally Intelligent Phone Agents

- 构建包含26个应用的持久化用户身份系统,支持跨应用数据联动
- 多任务测试中整体成功率52%,跨应用任务仅37%
- 适合研究个人化智能体与移动设备交互的开发者与研究人员
一个实用的手机助手需要具备个性化智能,能够基于设备上的用户身份、历史和偏好进行推理,而非仅在孤立环境中执行指令。现有移动端智能体评测基准缺乏这种个性化特征。我们提出iOSWorld,首个基于原生iOS的交互式仿真评测基准,围绕26个新开发的iOS应用构建,涵盖交易、消息、出行记录、社交关系及金融活动等连通数据。该基准包含133项任务,分为三类:单应用任务(27项)、多应用任务(60项)以及需从个人数据中推断模式的记忆与个性化任务(46项)。我们在视觉仅输入与特权视觉+XML两种设置下评估前沿及开源模型。最佳配置总体准确率达52%,但跨应用任务仅为37%。特权信息可使前沿模型提升最高26个百分点,而小型模型未从中受益。所有应用、种子数据、任务、评分标准与评估代码均已开源。
原文摘要 · Abstract (English)
A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated instructions in an impersonal sandbox. Existing mobile agent benchmarks lack this kind of personalization. We introduce iOSWorld, the first interactive native iOS simulator benchmark built around a persistent user identity spanning 26 newly built iOS apps. These apps contain connected data such as transactions, messages, travel records, social relationships, and financial activity. iOSWorld includes 133 tasks across three increasingly difficult categories. Single-app tasks (27) test one app, multi-app tasks (60) span 2 to 8 apps, and memory and personalization tasks (46) require agents to infer patterns from personal data. We evaluate frontier and open-source computer-use models in both vision-only and privileged vision+XML settings. The best configuration reaches 52\% overall but only 37\% on multi-app tasks. Privileged vision+XML access improves frontier models by up to 26 percentage points, while smaller models do not benefit from added accessibility-tree input. We release iOSWorld as an open-source benchmark with all apps, seeded data, tasks, rubrics, and evaluation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。