arXiv:2409.00063cs.CYcs.CL2024-09被引 25

用大模型生成城市出行数据,低成本替代传统调查。

Urban Mobility Assessment Using LLMs

  • 用大模型提示生成合成出行数据,避免隐私与成本问题。
  • 微调后的Llama-2生成数据在多层级上接近真实调查结果。
  • 适合城市规划、交通研究者快速获取出行模式数据。

理解城市出行模式并分析人们如何在城市中移动,有助于提升生活质量,并推动更宜居、高效和可持续的城市发展。然而,通过用户追踪或出行调查收集出行数据面临隐私顾虑、参与度低和成本高等挑战。本文提出一种基于AI的创新方法,通过提示大型语言模型(LLMs)合成出行调查数据,利用其丰富的背景知识和文本生成能力。研究在多个美国都会区评估该方法的有效性,对比合成数据与真实调查数据在不同粒度下的表现:(i) 模式层面,比较平均出行地点数和出行时间等聚合指标;(ii) 行程层面,通过转移概率对比完整行程;(iii) 活动链层面,分析个体访问地点的序列。实验涵盖若干专有和开源的LLM,发现如Llama-2等开源基础模型,仅需少量真实数据微调,即可生成高度接近真实调查数据的合成数据,为出行研究中使用此类数据提供了有力依据。

原文摘要 · Abstract (English)

Understanding urban mobility patterns and analyzing how people move around cities helps improve the overall quality of life and supports the development of more livable, efficient, and sustainable urban areas. A challenging aspect of this work is the collection of mobility data by means of user tracking or travel surveys, given the associated privacy concerns, noncompliance, and high cost. This work proposes an innovative AI-based approach for synthesizing travel surveys by prompting large language models (LLMs), aiming to leverage their vast amount of relevant background knowledge and text generation capabilities. Our study evaluates the effectiveness of this approach across various U.S. metropolitan areas by comparing the results against existing survey data at different granularity levels. These levels include (i) pattern level, which compares aggregated metrics like the average number of locations traveled and travel time, (ii) trip level, which focuses on comparing trips as whole units using transition probabilities, and (iii) activity chain level, which examines the sequence of locations visited by individuals. Our work covers several proprietary and open-source LLMs, revealing that open-source base models like Llama-2, when fine-tuned on even a limited amount of actual data, can generate synthetic data that closely mimics the actual travel survey data, and as such provides an argument for using such data in mobility studies.

城市出行大模型数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。