学习用户偏好分布,让决策更适应变化的上下文环境。
Contextual Preference Distribution Learning
- 用上下文特征预测偏好分布,而非固定点估计
- 在模拟拼车场景中,降低114倍的决策意外度
- 适合需要应对不确定偏好的风险敏感决策场景
决策问题常因异质且依赖上下文的人类偏好而存在不确定性。为此,我们提出一种顺序学习与优化框架,用于学习偏好分布并应用于下游问题,如风险规避型决策。研究聚焦可建模为(整数)线性规划的人类选择场景。现有逆向优化和选择建模方法虽能从观测选择中推断偏好,但通常仅生成点估计或无法捕捉上下文变化,难以支持风险规避决策。我们采用有界方差得分函数梯度估计器,训练一个将上下文特征映射到可参数化分布族的预测模型,获得最大似然估计。该模型在后续优化阶段生成未见上下文下的情景。在模拟拼车环境中,该方法相比具备完美预测的风险中性方案,平均事后意外度降低达114倍;相比领先的风险规避基线,降低达25倍。
原文摘要 · Abstract (English)
Decision-making problems often feature uncertainty stemming from heterogeneous and context-dependent human preferences. To address this, we propose a sequential learning-and-optimization pipeline to learn preference distributions and leverage them to solve downstream problems, for example risk-averse formulations. We focus on human choice settings that can be formulated as (integer) linear programs. In such settings, existing inverse optimization and choice modelling methods infer preferences from observed choices but typically produce point estimates or fail to capture contextual shifts, making them unsuitable for risk-averse decision-making. Using a bounded-variance score function gradient estimator, we train a predictive model mapping contextual features to a rich class of parameterizable distributions. This approach yields a maximum likelihood estimate. The model generates scenarios for unseen contexts in the subsequent optimization phase. In a synthetic ridesharing environment, our approach reduces average post-decision surprise by up to 114$\times$ compared to a risk-neutral approach with perfect predictions and up to 25$\times$ compared to leading risk-averse baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。