机器学习能更好捕捉复杂决策中的个体偏好,助力政策制定。
An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks
- 用四种机器学习模型对比学习五类行为决策规则。
- 非参数模型在多数场景下预测准确率提升6%至96%。
- 实际能源政策数据中,双神经网络表现最优,适合政策研究者。
离散选择建模常用于政策制定中的偏好收集,传统上依赖参数化模型。机器学习可通过数据驱动方式突破这一局限,学习个体偏好。然而,现有研究缺乏对机器学习在个体异质性条件下估计个体选择规则能力的系统评估,尤其是在偏好收集中常见挑战背景下。本研究评估了四种机器学习模型(多元逻辑回归、广义加性模型、双神经网络、高斯过程)在学习五种重要行为与社会科学决策规则(线性强效用、单调强效用、理想点、词典式半序、多属性线性球形累加器)方面的表现。通过蒙特卡洛实验,考察了三类因素的影响:选择方案属性数量增加、训练选择集数量增加、选择规则确定性增强。结果表明,半参数与非参数模型在所有规则与实验条件下普遍优于参数模型;模型性能随训练选择集数量增加提升6%至96%,随规则确定性增强提升0%至55%。一项基于真实能源政策偏好数据的案例研究显示,双神经网络(TNN)表现最佳,贝叶斯信息准则(BIC)为13.351。研究验证了半参数与非参数模型在政策导向离散选择建模中的可行性与局限性,并强调应根据任务情境选择合适模型。
原文摘要 · Abstract (English)
Discrete choice modeling is a common tool used for preference elicitation during policy-making, but this is typically done through parametric models. Machine learning can push the boundaries of discrete choice modeling for policy-based preference elicitation by adopting a data-driven approach or learning individual preferences. However, there is limited knowledge of how well machine learning methods can estimate individual discrete choice rules under individual heterogeneity, especially in the context of challenges often experienced during preference elicitation. This study evaluates four machine learning models (multinomial logistic regression, generalized additive model, twinned neural network, and Gaussian process) with respect to their capacity to learn and predict five choice rules that are important in the behavioral and social sciences (linear strong utility, monotonic strong utility, ideal point, lexicographic semiorder, and multiattribute linear ballistic accumulator). Monte Carlo experiments were performed to assess model performance when increasing a) the number of attributes in the choice alternatives, b) the number of training choice sets, and c) the choice rule's determinism. The simulation results demonstrated that semi-parametric and non-parametric models generally outperform parametric models across all choice rules and experimental contexts. Model performance also generally improves by 6% to 96% and 0% to 55%, respectively, with an increase in training choice sets and choice rule determinism. A case study using real energy policy preference data was also conducted, where TNN performed best with a BIC of 13.351. This work demonstrated the viability and limitations of semi-parametric and non-parametric models in the context of policy-centric discrete choice modeling and showed how the choice task context should drive model selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。