用社交媒体数据评测大模型的地理规划能力,发现其在复杂任务中表现骤降。
LifePlanner: Evaluating LLM Agents for Geo-spatial Planning with Social Media Data

- 引入社交媒体数据增强地图信息,构建多模态规划评估基准
- 复杂规划任务通过率仅40.2%,主要因证据获取不全和约束融合弱
- 适合研究具身智能、多模态推理与真实场景规划的学者参考
地理空间规划(如旅行路线设计)是检验大语言模型代理能力的真实场景,需结合工具使用、噪声信息检索及多约束推理。现有基准大多仅提供结构化地理数据,忽略了日常规划中常用的开放社交信号。本文提出LifePlanner,将大规模本地社交媒体帖子融入地图数据,并通过MCP工具集提供访问。该基准涵盖四大任务类别与三个难度层级。实验表明,前沿大模型在简单检索任务中表现良好,但在复杂规划中性能急剧下降,通过率降至40.2%。失败主因在于从大规模多模态数据库中获取不完整证据、工具使用不精确以及约束整合能力弱,而非模型规模或推理长度。结果表明,未来进展应聚焦有效具身规划,而非单纯模型扩展。
原文摘要 · Abstract (English)
Geo-spatial planning, like trip design, is a realistic testbed for LLM agents because it requires grounded tool use, noisy evidence retrieval, and multi-constraint reasoning. Most benchmarks, however, only provide clean geospatial data and tools, missing the open-ended social signals that people use in daily planning. We introduce LifePlanner, a benchmark that enriches map data with large-scale local social media posts and provides access through an MCP toolset. LifePlanner provides an evaluation suite spanning four task categories and three difficulty levels. Experiments show frontier LLMs perform well on simple retrieval but degrade sharply on complex planning, with the Pass Rate dropping to 40.2%. Results show that failures mainly stem from incomplete evidence acquisition from such a large multimodal database, imprecise tool use, and weak constraint integration rather than model size or reasoning length, suggesting that future progress requires effective grounded planning instead of scaling alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。