arXiv:2503.12761cs.AI2025-03被引 3

用可解释深度逆强化学习分析出行行为决策机制。

Analyzing sequential activity and travel decisions with interpretable deep inverse reinforcement learning

  • 基于对抗性逆强化学习推断出行行为的奖励与策略函数。
  • 揭示不同人群出行选择的关键影响因素及偏好模式。
  • 适合交通规划、行为建模研究者参考使用。

出行需求建模正从传统的以行程为中心的模型转向以行为为导向的活动为基础模型,因为日常出行本质上由人类活动驱动。为分析连续的活动-出行决策,深度逆强化学习(DIRL)已被证明可通过深度神经网络近似表示偏好的奖励函数和复制观测行为的策略函数,有效学习决策机制。然而,现有研究大多聚焦提升预测准确性,较少关注解释序列决策背后的内在机制。为此,本文提出一种可解释的DIRL框架,弥合数据驱动机器学习与理论驱动行为模型之间的鸿沟。该框架采用对抗性逆强化学习方法,推断活动-出行行为的奖励与策略函数。通过基于策略函数中选择概率的代理可解释模型来解读策略函数,同时通过推导不同活动-出行模式的短期奖励与长期回报来解释奖励函数。对真实出行调查数据的分析揭示了两大关键成果:(i) 策略函数提供的行为模式洞察,突出决策中的关键因素及社会人口群体间的差异;(ii) 奖励函数提供的行为偏好洞察,表明个体从特定活动序列中获得的效用。

原文摘要 · Abstract (English)

Travel demand modeling has shifted from aggregated trip-based models to behavior-oriented activity-based models because daily trips are essentially driven by human activities. To analyze the sequential activity-travel decisions, deep inverse reinforcement learning (DIRL) has proven effective in learning the decision mechanisms by approximating a reward function to represent preferences and a policy function to replicate observed behavior using deep neural networks (DNNs). However, most existing research has focused on using DIRL to enhance only prediction accuracy, with limited exploration into interpreting the underlying decision mechanisms guiding sequential decision-making. To address this gap, we introduce an interpretable DIRL framework for analyzing activity-travel decision processes, bridging the gap between data-driven machine learning and theory-driven behavioral models. Our proposed framework adapts an adversarial IRL approach to infer the reward and policy functions of activity-travel behavior. The policy function is interpreted through a surrogate interpretable model based on choice probabilities from the policy function, while the reward function is interpreted by deriving both short-term rewards and long-term returns for various activity-travel patterns. Our analysis of real-world travel survey data reveals promising results in two key areas: (i) behavioral pattern insights from the policy function, highlighting critical factors in decision-making and variations among socio-demographic groups, and (ii) behavioral preference insights from the reward function, indicating the utility individuals gain from specific activity sequences.

行为建模逆强化学习出行决策可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。