提出一套决策支持系统,判断市场策略能否上线,避免盲目部署。
Decision Support for Marketplace Policies under Incomplete Evidence: From Replay to Launch Readiness

- 构建可复现的决策支持框架,融合回放与多维验证
- 实测新策略在回放中收益提升47.7%,保守下界仍增45.8%
- 适合需要安全上线决策的平台方,尤其关注因果效应的场景
市场平台常通过历史数据评估定价与分配策略,但离线表现优异并不意味着可直接部署。在实时竞价(RTB)市场中,底价和门槛策略的变化不仅影响收入,还牵动成交率、广告主价值、预算进度及拍卖间竞争,引发反馈与干扰。核心问题并非评估策略是否提升离线指标,而是判断现有证据是否足以支持直接上线或仅需进一步验证。为此,我们提出一种支持感知的决策支持系统(DSS),区分有潜力与可行动的证据。该框架整合回放、支持感知的离线策略评估(OPE)、保守下界排序、多方约束、跨时验证、敏感性分析及干扰感知验证设计,形成保持主张的可审计流水线,输出上线就绪分类而非单一性能估计。基于iPinYou风格的RTB日志应用该框架,识别出一种边际门控底价策略为最优候选,其回放收益提升47.7%,保守下尾提升45.8%,且跨时表现稳定。然而,系统不建议直接上线。消融实验表明,简化流程虽选出同一策略却错误推荐部署,未解决倾向性、竞标者响应与干扰等关键因果假设。相比之下,本DSS选择相同策略但转为在线验证,揭示了证据缺失。总体贡献是提供一个可复现的DSS协议,防止部分识别下的决策过度主张,并将离线评估转化为可审计、可行动的推荐。
原文摘要 · Abstract (English)
Marketplace platforms routinely evaluate pricing and allocation policies using logged observational data, yet strong offline performance does not imply that a policy is safe to deploy. In real-time bidding (RTB) marketplaces, reserve-price and floor-policy changes affect not only revenue but also fill, advertiser value, budget pacing, and competition across auctions, creating feedback and interference. The central problem is therefore not to estimate whether a policy improves an offline metric, but to determine whether the available evidence justifies direct launch or only further validation. In this regard, we propose a support-aware decision-support system (DSS) that distinguishes promising from actionable evidence. The framework integrates replay, support-aware off-policy evaluation (OPE), conservative lower-bound ranking, multi-sided guardrails, out-of-time validation, sensitivity analysis, and interference-aware validation design into a claim-preserving pipeline that outputs a launch-readiness classification rather than a single performance estimate. Applying the framework to iPinYou-style RTB logs, we identify a margin-gated floor policy as the leading candidate, with a 47.7% replay yield lift, a 45.8% conservative lower-tail lift, and stable out-of-time performance. However, the framework does not recommend direct launch. A decision-rule ablation shows that simplified pipelines select the same policy but incorrectly recommend deployment, leaving key causal assumptions unresolved. In contrast, the proposed DSS selects the same policy but changes the action to online validation, reflecting missing evidence on propensities, bidder response, and interference. Overall, the contribution is a reproducible DSS protocol that prevents decision overclaim under partial identification and converts offline evaluation into an auditable, action-oriented recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。