arXiv:2605.00193cs.LGstat.ML2026-05

提出OTSS模型,让决策权重自动适配上下文环境。

OTSS: Output-Targeted Soft Segmentation for Contextual Decision-Weight Learning

论文配图:OTSS: Output-Targeted Soft Segmentation for Contextual Decision-Weight Learning
图 1 · 摘自论文原文
  • 用软分段方法学习可解释因子的个性化权重向量
  • 在重叠场景下均方误差最低,比最强基线快100倍
  • 适合需要动态调整决策权重的推荐与营销系统

许多机器学习系统通过优化分解目标做出受限决策,但上下文相关的优化目标常被视为固定。本文研究上下文决策权重学习:从记录的决策和代理输出中,学习一个面向优化器的权重向量 w(x),作用于可解释的决策因子 z(x,d),而非直接策略或通用预测得分。提出 OTSS 模型——一种面向输出的目标软分段方法,部署个性化决策就绪权重向量。理论层面揭示硬划分与软划分的本质差异:硬划分在重叠情形下存在近似-估计权衡,而真实可实现的固定-K 软类能消除硬划分的近似下限,达到参数化速率。在有限评估库的受控基准测试中,真权重向量与下游遗憾可精确计算。在典型重叠设置中,OTSS 的平均遗憾低于所有对比方法,包括最强的软混合基线 EM 混合回归;其系数恢复表现与 EM 相当,但速度提升约两个数量级。在 K=5 的匹配基准中,即使在硬路由真实条件下仍具竞争力,并随异质性减弱与样本量增加而持续改进。在包含真实家庭协变量与动作结构的 Complete Journey 零售锚点上,OTSS 再次取得最低平均遗憾点估计。

原文摘要 · Abstract (English)

Many machine learning systems make constrained decisions by optimizing factorized objectives, but the context-specific objective is often treated as fixed. We study contextual decision-weight learning: from logged decisions and proxy outputs, learn an optimizer-facing weight vector w(x) over interpretable decision factors z(x,d), rather than a direct policy or generic predictive score. We propose OTSS, an output-targeted soft-segmentation model that deploys the personalized decision-ready weight vector. At the function-class level, the theory highlights a hard-versus-soft distinction. Hard partitions incur an approximation-estimation tradeoff under overlap, while a realizable fixed-K soft class removes the hard-partition approximation floor and attains a parametric rate. We evaluate OTSS in controlled benchmarks with finite evaluation libraries, where the true weight vector and downstream regret can be computed exactly. In the representative overlap setting, OTSS attains the lowest mean regret among the comparators, including EM mixture regression, the strongest soft-mixture baseline in our comparison; it matches EM on coefficient recovery while running about two orders of magnitude faster. In a matched K=5 benchmark, OTSS remains competitive under hard-routed truth and improves as heterogeneity becomes softer and sample size grows. On a fixed Complete Journey retail anchor with real household covariates and action geometry, OTSS again achieves the lowest mean-regret point estimate.

决策学习软分段权重优化上下文适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。