用少量标注数据训练弱模型,再用其生成标签提升强模型决策能力。
Weak-to-Strong Learning in Decision Making

- 先用少样本训练弱模型,再用它给无标签数据生成预测分布作为软标签。
- 理论证明:当特征相关维度低时,大量无标签数据能降低教师错误影响。
- 适合数据标注成本高、需优化决策的场景,如供应链、内容审核。
许多实际决策依赖于在可观测上下文条件下估计不确定结果的预测模型。然而,训练这类模型常面临根本性的数据不对称:标注结果稀缺或获取成本高,而上下文协变量却丰富。针对这一问题,我们提出一种决策感知的弱到强(W2S)框架,利用有标签和无标签数据共同提升上下文随机优化性能。首先,基于有限有标签数据训练一个弱模型;随后,用该模型为无标签上下文生成预测结果分布,作为强模型训练的软监督信号。我们建立了W2S的超决策风险非渐近上界,并给出了强模型仅依赖有标签数据的互补下界。两者的比较揭示了在何种条件下W2S可提升下游决策表现。关键指标是弱模型与强模型特征表示之间的相关维数:当其较小时,大量无标签数据能有效抑制教师误差在非重叠方向上的影响。合成的新刊商实验与基于真实数据的评论审核实验均验证了理论预期。
原文摘要 · Abstract (English)
Many operational decisions rely on predictive models that estimate uncertain outcomes conditional on observable contexts. Training such models, however, often faces a fundamental data asymmetry: labeled outcomes are scarce or costly to obtain, while contextual covariates are abundant. Motivated by this data asymmetry, we develop a decision-aware weak-to-strong (W2S) framework that leverages both labeled and unlabeled data to improve contextual stochastic optimization. Specifically, we first train a weak model using limited labeled data and then use it to generate predicted outcome distributions on unlabeled contexts. These distributions provide soft supervision for training a strong model. We establish a non-asymptotic upper bound on the excess decision risk of W2S and a complementary lower bound for a strong-only benchmark. Their comparison yields explicit sufficient conditions under which W2S improves downstream decision performance. The key quantity is the correlation dimension between the weak and strong feature representations: when it is small, abundant unlabeled data reduce the effect of teacher errors along non-overlapping directions. A synthetic newsvendor experiment and a comment moderation experiment based on real-world data provide empirical evidence consistent with the theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。