arXiv:2608.03018cs.AI2026-08

让智能体跨系统完成城市任务,准确率比现有方法高10个百分点。

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

论文配图:UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
图 1 · 摘自论文原文
  • 用大模型+工具链实现自然语言到跨系统操作的转化。
  • 在城市任务中达成71%成功率,显著优于基线模型。
  • 专为城市服务设计评估基准,兼顾执行质量与结果准确性。

现代城市依赖越来越多的数字服务运行,但居民日常需求仍难以满足。服务碎片化、缺乏互操作性,给用户带来沉重操作负担。现有数字平台、城市基础模型和智能助手仅解决城市任务的局部问题,难以可靠地将复杂自然语言请求转化为可执行的跨系统工作流。本文提出Urban-Agent,一种面向跨系统城市任务的工具增强型智能体框架。它结合大语言模型的认知与推理能力,以及支持代码执行、API调用和模型上下文协议的工具集。通过一个自适应闭环,它能在行动前澄清缺失信息,基于实时观测使用工具,并确保最终响应与观察证据及任务约束一致。为弥补评估空白,我们引入Urban-Eval,一个专门针对跨系统城市请求设计的基准。该基准同时评估任务结果与执行质量,包括所需工具覆盖率、依赖关系有效性与证据可追溯性。实验表明,Urban-Agent在GPT-5-mini、Gemini-2.5-flash、DeepSeek-V4-flash和Qwen3-235B-A22B上均达到71%的任务成功率,比最强基线高出10个百分点。

原文摘要 · Abstract (English)

Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.

智能体城市服务工具增强任务自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。