用伪样本匹配融合小而偏的实验数据,提升网约车定价模型效果
Augmenting Limited and Biased RCTs through Pseudo-Sample Matching-Based Observational Data Fusion Method
- 从低质量实验数据生成伪样本,与真实观测数据匹配扩充数据集
- 线上实验提升0.41%利润,百万级营收场景下收益显著
- 解决工业场景中实验数据少、偏差大、高维特征难处理问题
在网约车定价场景中,企业常通过随机对照试验(RCT)和提升模型评估折扣对订单的影响,这对市场竞争结果有重要影响。但受成本限制,实验数据占总流量仅0.65%,且工业流程复杂导致数据存在异质性、干扰和选择偏差,难以修正。现有数据融合方法在高维特征和现实数据条件下难以有效实施。为此,本文提出一种基于伪样本匹配的实证数据融合方法:从有偏的低质量RCT数据生成伪样本,并与大规模观测数据中最相似的样本匹配,以扩充实验数据并缓解异质性。通过仿真实验及真实数据的离线与在线测试验证,一周在线实验实现0.41%利润提升,在亿级营收场景中具有显著价值。同时指出低质量实验数据对模型训练、离线评估和在线收益的损害,强调提升工业场景中实验数据质量的重要性。更多仿真细节见GitHub仓库:https://github.com/Kairong-Han/Pseudo-Matching。
原文摘要 · Abstract (English)
In the online ride-hailing pricing context, companies often conduct randomized controlled trials (RCTs) and utilize uplift models to assess the effect of discounts on customer orders, which substantially influences competitive market outcomes. However, due to the high cost of RCTs, the proportion of trial data relative to observational data is small, which only accounts for 0.65\% of total traffic in our context, resulting in significant bias when generalizing to the broader user base. Additionally, the complexity of industrial processes reduces the quality of RCT data, which is often subject to heterogeneity from potential interference and selection bias, making it difficult to correct. Moreover, existing data fusion methods are challenging to implement effectively in complex industrial settings due to the high dimensionality of features and the strict assumptions that are hard to verify with real-world data. To address these issues, we propose an empirical data fusion method called pseudo-sample matching. By generating pseudo-samples from biased, low-quality RCT data and matching them with the most similar samples from large-scale observational data, the method expands the RCT dataset while mitigating its heterogeneity. We validated the method through simulation experiments, conducted offline and online tests using real-world data. In a week-long online experiment, we achieved a 0.41\% improvement in profit, which is a considerable gain when scaled to industrial scenarios with hundreds of millions in revenue. In addition, we discuss the harm to model training, offline evaluation, and online economic benefits when the RCT data quality is not high, and emphasize the importance of improving RCT data quality in industrial scenarios. Further details of the simulation experiments can be found in the GitHub repository https://github.com/Kairong-Han/Pseudo-Matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。