arXiv:2605.10782cs.AI2026-05被引 1

构建多任务城市轨迹语言对齐基准,打通轨迹与语义描述的桥梁。

TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding

论文配图:TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding
图 1 · 摘自论文原文
  • 提出三任务统一框架:指令生成轨迹、语义检索轨迹、轨迹文本描述。
  • 涵盖30万条真实城市轨迹,产生210万任务实例,覆盖三座城市。
  • 首次实现轨迹与语言在细粒度上的可验证对齐,适合多模态研究者使用。

城市出行既表现为空间轨迹,也通过自然语言描述出行意图、约束和偏好。然而,以往研究很少在同一组真实轨迹上联合评估这两种模态:轨迹建模多以几何为中心,而语言导向的基准则侧重路线规划或工具使用,缺乏对文本与底层轨迹间细粒度、可验证对齐的评估。本文提出TrajPrism,一个面向语言-轨迹对齐的多任务基准,包含(i)指令条件下的轨迹生成,(ii)语言驱动的语义轨迹检索,以及(iii)轨迹描述生成,并配套评估协议,衡量轨迹保真度、检索质量与语言依存性。我们基于四个维度的出行意图分类体系,对真实城市轨迹进行标注,构建了包含波尔图、旧金山和北京共30万条精选轨迹的基准,生成210万任务实例(三种指令变体、三种检索查询、每条轨迹一个描述)。我们进一步开发了三个概念验证模型:TrajAnchor(指令轨迹生成)、TrajFuse(语义轨迹检索)、TrajRap(轨迹描述生成)。这些模型展示了仅依赖几何的基线在新协议下表现显著不足,尤其当语言作为输入输出接口时差距更大。我们开源了TrajPrism代码与可复现的标注流程,支持在具备轨迹与地图资源的城市间迁移。

原文摘要 · Abstract (English)

Urban mobility is naturally expressed both as trajectories in space and as natural-language descriptions of travel intent, constraints, and preferences. However, prior work rarely evaluates these two modalities together on the same real-world trajectories: trajectory modeling often stays geometry-centric, while language-centric mobility benchmarks frequently target route planning and tool use rather than fine-grained, verifiable alignment between text and the underlying route. We introduce TrajPrism, a multi-task benchmark for language-trajectory alignment that unifies (i) instruction-conditioned trajectory generation, (ii) language-driven semantic trajectory retrieval, and (iii) trajectory captioning, together with an evaluation protocol that measures trajectory fidelity, retrieval quality, and language groundedness. We construct TrajPrism by pairing real urban trajectories with judge-filtered language annotations generated under a four-dimensional travel-intent taxonomy. The benchmark contains 300K selected trajectories across Porto, San Francisco, and Beijing, yielding 2.1M task instances from three instruction variants, three retrieval queries, and one caption per trajectory. We further develop proof-of-concept models for each task: TrajAnchor for instruction-conditioned trajectory generation, TrajFuse for semantic trajectory retrieval, and TrajRap for trajectory captioning. These models instantiate the proposed tasks and show that geometry-only trajectory baselines leave a large gap on our protocol, especially where language is part of the input-output interface. We release TrajPrism with code and a reproducible annotation pipeline that is designed to be portable across cities, given compatible trajectory inputs and map resources.

轨迹理解多模态语言对齐城市计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。