arXiv:2506.06410econ.GNcs.LG2025-06被引 1

用强化学习自动设计选择模型,减少人工试错。

Delphos: A reinforcement learning framework for assisting discrete choice model specification

  • 将模型构建转化为序列决策问题,智能体逐步选择变量、变换等操作。
  • 在模拟和真实数据上生成性能优异且符合行为预期的模型,仅探索小部分可行空间。
  • 适合需要高效建模的交通、经济等领域研究者,降低建模门槛。

我们提出Delphos,一种基于深度强化学习的离散选择模型构建辅助框架。该框架将模型规范过程视为序列决策问题,受人类建模者逐步推理方式启发。智能体通过选择变量、设定通用或特定于选项的效用参数、应用非线性变换及包含协变量交互等建模动作,在与建模环境交互中学习生成高性能候选模型。系统采用Deep Q-Network,依据对数似然等建模结果和参数符号等行为预期,接收延迟奖励信号,并将奖励分布至整个动作序列,以识别有效建模决策。我们在模拟和真实数据集上验证了该方法,结果显示智能体能自适应探索策略,在仅覆盖少量可行建模空间的前提下,生成性能优良且行为合理的模型。实验表明,Delphos具备辅助建模过程的潜力,为自动化模型构建提供了新思路。

原文摘要 · Abstract (English)

We introduce Delphos, a deep reinforcement learning framework for assisting the discrete choice model specification process. Delphos aims to support the modeller by providing automated, data-driven suggestions for utility specifications, thereby reducing the effort required to develop and refine utility functions. Delphos conceptualises model specification as a sequential decision-making problem, inspired by the way human choice modellers iteratively construct models through a series of reasoned specification decisions. In this setting, an agent learns to specify high-performing candidate models by choosing a sequence of modelling actions, such as selecting variables, accommodating both generic and alternative-specific taste parameters, applying non-linear transformations, and including interactions with covariates, while interacting with a modelling environment that estimates each candidate and returns a reward signal. Specifically, Delphos uses a Deep Q-Network that receives delayed rewards based on modelling outcomes (e.g., log-likelihood) and behavioural expectations (e.g., parameter signs), and distributes this signal across the sequence of actions to learn which modelling decisions lead to well-performing candidates. We evaluate Delphos on both simulated and empirical datasets using multiple reward settings. In simulated cases, learning curves, Q-value patterns, and performance metrics show that the agent learns to adaptively explore strategies to propose well-performing models across search spaces, while covering only a small fraction of the feasible modelling space. We further apply the framework to two empirical datasets to demonstrate its practical use. These experiments illustrate the ability of Delphos to generate competitive, behaviourally plausible models and highlight the potential of this adaptive, learning-based framework to assist the model specification process.

强化学习模型构建选择模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。