arXiv:2509.18047stat.MLcs.LG2025-09

用机器学习建模个体偏好差异,提升面板数据预测精度。

Functional effects models: Accounting for preference heterogeneity in panel data with machine learning

  • 用梯度提升树和神经网络学习人口特征对偏好的影响函数
  • 在小样本下仍能准确预测未见个体的选择行为
  • 适合需要个性化建模的推荐系统与市场研究场景

本文提出一种通用的函数效应模型(Functional Effects Model),利用机器学习方法从社会人口特征中学习个体特异性偏好参数,从而在面板选择数据中捕捉个体间异质性。该模型相比传统固定效应和随机/混合效应模型有三大优势:(i)将个体效应表示为人口变量的函数,可预测未曾观察到的个体选择;(ii)通过近似最大似然估计规避固定效应模型的偶然参数问题,即使每人观测次数较少也有效;(iii)无需随机效应模型强分布假设,更贴近真实情况。采用梯度提升决策树和深度神经网络等非线性回归器学习函数截距与斜率。在合成实验及三个真实世界面板案例研究中验证:(i)当数据生成过程已知时,可准确识别个体效应真实值;(ii)在预测性能上优于忽略异质性的先进机器学习选择模型,且在学习个体差异方面超越传统静态面板模型。结果表明,将函数效应模型的个体常数与RUMBoost的复杂非线性效用结合的FI-RUMBoost模型,在大规模显性偏好面板数据上表现最优。

原文摘要 · Abstract (English)

In this paper, we present a general specification for Functional Effects Models, which use Machine Learning (ML) methodologies to learn individual-specific preference parameters from socio-demographic characteristics, therefore accounting for inter-individual heterogeneity in panel choice data. We identify three specific advantages of the Functional Effects Model over traditional fixed, and random/mixed effects models: (i) by mapping individual-specific effects as a function of socio-demographic variables, we can account for these effects when forecasting choices of previously unobserved individuals (ii) the (approximate) maximum-likelihood estimation of functional effects avoids the incidental parameters problem of the fixed effects model, even when the number of observed choices per individual is small; and (iii) we do not rely on the strong distributional assumptions of the random effects model, which may not match reality. We learn functional intercept and functional slopes with powerful non-linear machine learning regressors for tabular data, namely gradient boosting decision trees and deep neural networks. We validate our proposed methodology on a synthetic experiment and three real-world panel case studies, demonstrating that the Functional Effects Model: (i) can identify the true values of individual-specific effects when the data generation process is known; (ii) outperforms both state-of-the-art ML choice modelling techniques that omit individual heterogeneity in terms of predictive performance, as well as traditional static panel choice models in terms of learning inter-individual heterogeneity. The results indicate that the FI-RUMBoost model, which combines the individual-specific constants of the Functional Effects Model with the complex, non-linear utilities of RUMBoost, performs marginally best on large-scale revealed preference panel data.

机器学习面板数据偏好建模异质性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。