arXiv:2512.19234cs.AI2025-12被引 5

用真实外卖场景测试AI Agent的长期赚钱能力,发现它们短视且常违规。

DeliveryBench: Can Agents Earn Profit in Real World?

  • 在生成的3D城市中模拟外卖员长时间接单赚钱任务
  • 多模型表现远低于人类,普遍忽略配送时限和资源消耗
  • 揭示不同模型性格差异,适合研究具身智能与约束决策

大语言模型和视觉语言模型正越来越多地作为具身智能体部署,但现有基准多局限于简单短期任务,难以反映真实世界决策中的复杂约束。为弥补这一差距,我们提出DeliveryBench,一个基于真实外卖职业的城市级具身基准。外卖员需在数小时长周期内最大化净收益,同时应对配送时限、交通成本、车辆电量及与其他骑手、顾客的交互等多重约束。DeliveryBench在程序生成的3D城市中构建了多样道路网络、建筑布局、功能地点、交通方式与真实资源动态,支持对约束感知、长周期规划能力的系统评估。我们在九个城市对多种视觉语言模型驱动的智能体进行评测,并与人类玩家对比。结果表明,智能体与人类存在显著性能差距,普遍存在短视行为,频繁违反基本常识约束。此外,观察到模型间呈现明显性格差异(如冒险型GPT-5 vs. 保守型Claude),凸显当前视觉语言模型在高约束现实环境中的脆弱性与多样性。代码、数据与基准已公开于https://deliverybench.github.io。

原文摘要 · Abstract (English)

LLMs and VLMs are increasingly deployed as embodied agents, yet existing benchmarks largely revolve around simple short-term tasks and struggle to capture rich realistic constraints that shape real-world decision making. To close this gap, we propose DeliveryBench, a city-scale embodied benchmark grounded in the real-world profession of food delivery. Food couriers naturally operate under long-horizon objectives (maximizing net profit over hours) while managing diverse constraints, e.g., delivery deadline, transportation expense, vehicle battery, and necessary interactions with other couriers and customers. DeliveryBench instantiates this setting in procedurally generated 3D cities with diverse road networks, buildings, functional locations, transportation modes, and realistic resource dynamics, enabling systematic evaluation of constraint-aware, long-horizon planning. We benchmark a range of VLM-based agents across nine cities and compare them with human players. Our results reveal a substantial performance gap to humans, and find that these agents are short-sighted and frequently break basic commonsense constraints. Additionally, we observe distinct personalities across models (e.g., adventurous GPT-5 vs. conservative Claude), highlighting both the brittleness and the diversity of current VLM-based embodied agents in realistic, constraint-dense environments. Our code, data, and benchmark are available at https://deliverybench.github.io.

具身智能长周期规划约束决策真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。