arXiv:2512.14713cs.LGecon.EM2025-12

用贝叶斯潜类强化学习建模司机如何根据反馈调整出行偏好。

A Bayesian latent class reinforcement learning framework to capture adaptive, feedback-driven travel behaviour

  • 基于变分贝叶斯的潜类强化学习框架,捕捉个体偏好动态演化
  • 识别出三类不同适应策略:情境依赖型、持续占优型、探索结合情境型
  • 适用于研究个性化出行行为,尤其适合交通规划与智能导航系统

许多出行决策涉及经验形成过程,个体随时间学习自身偏好。同时,旅行者之间存在显著异质性,不仅体现在初始偏好上,也体现在偏好的演化方式上。本文提出一种潜类强化学习(LCRL)模型,以同时捕捉上述两种现象。我们利用驾驶模拟器数据集进行参数估计,采用变分贝叶斯方法。结果识别出三类显著不同的个体:第一类表现出情境依赖的偏好与情境特定的占优倾向;第二类无论情境如何均采取持续占优策略;第三类则结合探索行为与情境相关偏好。

原文摘要 · Abstract (English)

Many travel decisions involve a degree of experience formation, where individuals learn their preferences over time. At the same time, there is extensive scope for heterogeneity across individual travellers, both in their underlying preferences and in how these evolve. The present paper puts forward a Latent Class Reinforcement Learning (LCRL) model that allows analysts to capture both of these phenomena. We apply the model to a driving simulator dataset and estimate the parameters through Variational Bayes. We identify three distinct classes of individuals that differ markedly in how they adapt their preferences: the first displays context-dependent preferences with context-specific exploitative tendencies; the second follows a persistent exploitative strategy regardless of context; and the third engages in an exploratory strategy combined with context-specific preferences.

出行行为强化学习贝叶斯建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。