用语言模型模拟人群行为,仅凭统计数据就精准还原人群去向分布。
Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation

- 通过迭代比例调整,将语言模型的出行偏好对齐真实流量数据。
- 在两场棒球赛数据上,目的地分布误差降低25%,无需运行时修正。
- 适合隐私受限场景下的人群仿真,尤其适用于无轨迹数据的情况。
行人模拟器需为每个代理设定行为规则,但隐私限制通常只提供聚合统计数据,如区域设备数量和起讫点(OD)流量,缺乏个体轨迹信息。此类聚合数据无法唯一确定个体行为,因多种决策组合可产生相同统计结果。本文微调语言模型作为人群代理,使模拟群体的目的地构成匹配观测到的目的地组成(即各兴趣点的出发人群占比)。目标值来自OD流量数据,通过迭代比例拟合将模型自身的目的地分布重加权至该目标。由于微调会放大主导目的地类别,因此对重采样的轨迹使用低秩适配器,并校正训练组成,以在微调后达到目标分布。在两次棒球赛的移动网络统计数据上,微调后的代理无需运行时修正,目的地占比误差下降25%,网格相关性在不同策略间保持相似。
原文摘要 · Abstract (English)
Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for setting one to aggregate statistics, namely zone-level device counts and origin-to-destination (OD) flows, with no individual trajectories. Such aggregates under-determine individual behaviour, because many different sets of decisions reproduce the same counts. We fine-tune a language model crowd agent so that the simulated population matches the observed destination composition, the fraction of the departing crowd heading to each point of interest. We read this target from the OD flow and reweight the model's own destination distribution onto it by iterative proportional fitting. Because fine-tuning inflates the dominant destination class, we fit the low-rank adapter to trajectories resampled to a corrected training composition that reaches the target after this inflation. On mobile network counts from two baseball games the fine-tuned agent runs without inference-time correction, cutting the destination-share error by 25%, while the grid correlation remains similar across policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。