arXiv:2510.02333cs.CLcs.AI2025-10被引 1

用大模型生成社交媒体文本,丰富真实轨迹数据的语义信息。

Human Mobility Datasets Enriched With Contextual and Social Dimensions

  • 基于真实GPS轨迹,融合地点、交通方式、天气等上下文信息。
  • 首次引入大模型生成的虚构社交帖子,实现多模态轨迹分析。
  • 支持知识图谱构建,适合城市研究与智能系统开发人员使用。

本文介绍两个公开可用的语义增强型人类轨迹数据集及其构建流程。轨迹数据来自OpenStreetMap的公开GPS记录,覆盖巴黎和纽约两座结构迥异的大城市。每个数据集包含停靠点、移动段、兴趣点(POIs)、推断出的出行方式及天气数据等上下文层。创新性地引入由大语言模型(LLMs)生成的合成、真实感社交帖子,支持多模态与语义化移动行为分析。数据以表格和资源描述框架(RDF)格式提供,符合可发现、可访问、可互操作、可重用(FAIR)数据标准。我们提供的开源可复现管道支持数据定制,适用于行为建模、移动预测、知识图谱构建及基于LLM的应用研究。据我们所知,该资源是首个将真实移动数据、结构化语义增强、大模型生成文本与语义网络兼容性结合的可复用框架。

原文摘要 · Abstract (English)

In this resource paper, we present two publicly available datasets of semantically enriched human trajectories, together with the pipeline to build them. The trajectories are publicly available GPS traces retrieved from OpenStreetMap. Each dataset includes contextual layers such as stops, moves, points of interest (POIs), inferred transportation modes, and weather data. A novel semantic feature is the inclusion of synthetic, realistic social media posts generated by Large Language Models (LLMs), enabling multimodal and semantic mobility analysis. The datasets are available in both tabular and Resource Description Framework (RDF) formats, supporting semantic reasoning and FAIR data practices. They cover two structurally distinct, large cities: Paris and New York. Our open source reproducible pipeline allows for dataset customization, while the datasets support research tasks such as behavior modeling, mobility prediction, knowledge graph construction, and LLM-based applications. To our knowledge, our resource is the first to combine real-world movement, structured semantic enrichment, LLM-generated text, and semantic web compatibility in a reusable framework.

轨迹分析大模型城市计算语义增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。