arXiv:2511.06260cs.GTcs.AI2025-11被引 2

用代表性智能体降低交通建模计算成本,同时保持行为可解释性。

LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling

  • 为同质出行群体设计单一代表智能体,替代每个个体调用大模型。
  • 在经典与复杂场景中均快速收敛至用户均衡,动态稳定且可解释。
  • 适合研究交通行为、政策评估的学者与城市规划者参考。

大型语言模型(LLMs)被越来越多地用作基于代理的交通模型中自利出行者的行为代理。尽管相比传统模型更具灵活性和泛化能力,但因需为每位出行者调用一次大模型而面临可扩展性挑战。此外,大模型代理常做出不透明选择,导致日间动态不稳定。为此,我们提出以单个代表型大模型代理模拟具有相同决策情境的同质出行群体,其行为反映群体平均特征,并随时间更新路线混合策略,使其与群体总流量比例一致。每日,该代理回顾出行体验,标记期望更多使用的高奖励路径;随后通过可调节(渐进衰减)步长的可解释更新规则将其判断转化为策略调整。该设计提升可扩展性,分离推理与更新过程增强了决策逻辑清晰度并稳定学习。在经典交通分配场景中,方法迅速收敛至用户均衡;在包含收入异质性、多准则成本及多模式选择的复杂场景中,生成动态仍稳定可解释,复现了心理学与经济学中广泛记录的行为模式,如收费路与免费路选择中的诱饵效应,以及高收入者在驾车、公交与停车换乘间更愿为便利支付更高费用的现象。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as behavioral proxies for self-interested travelers in agent-based traffic models. Although more flexible and generalizable than conventional models, the practical use of these approaches remains limited by scalability due to the cost of calling one LLM for every traveler. Moreover, it has been found that LLM agents often make opaque choices and produce unstable day-to-day dynamics. To address these challenges, we propose to model each homogeneous traveler group facing the same decision context with a single representative LLM agent who behaves like the population's average, maintaining and updating a mixed strategy over routes that coincides with the group's aggregate flow proportions. Each day, the LLM reviews the travel experience and flags routes with positive reinforcement that they hope to use more often, and an interpretable update rule then converts this judgment into strategy adjustments using a tunable (progressively decaying) step size. The representative-agent design improves scalability, while the separation of reasoning from updating clarifies the decision logic while stabilizing learning. In classic traffic assignment settings, we find that the proposed approach converges rapidly to the user equilibrium. In richer settings with income heterogeneity, multi-criteria costs, and multi-modal choices, the generated dynamics remain stable and interpretable, reproducing plausible behavioral patterns well-documented in psychology and economics, for example, the decoy effect in toll versus non-toll road selection, and higher willingness-to-pay for convenience among higher-income travelers when choosing between driving, transit, and park-and-ride options.

交通建模大模型强化学习行为模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。