让合成数据真正支持营销决策,确保推荐结果与真实数据一致。
Trustworthy synthetic data for campaign decision support: strategy simulation fidelity and the PolicySynth framework
- 提出策略仿真保真度(SSF)衡量合成数据是否给出相同投放决策。
- PolicySynth框架通过条件生成使决策结构对齐,稳定性和准确性更高。
- 适合需要可信合成数据的金融、电信等行业的决策系统部署。
决策支持系统(DSS)越来越多地在合成客户群体上运行留存率的假设分析,因为隐私限制禁止直接使用真实数据。一个可信的系统必须确保合成数据能引导管理者做出与真实数据相同的决策;然而现有标准仅验证分布相似性,而非决策一致性,导致合成数据虽匹配所有边缘分布,仍可能误导营销团队选择错误的推广活动。本文提出三项贡献:策略仿真保真度(SSF),用于衡量合成数据与真实数据在“投放/不投放”决策上的一致频率;PolicySynth框架,其生成器基于生产环境的流失评分器进行条件建模,以对齐决策相关结构;以及由决策一致性、成员推断抵抗性和新记录率组成的三轴报告标准,作为部署最低质量门槛。在电信流失数据集和银行获客数据集上,PolicySynth的平均SSF分别达到0.923和0.960,种子间方差比CTGAN在电信领域低约10倍,在银行领域低2.5倍。其稳定性表现可落地:每月重训练时,投放建议变动不超过1.2个百分点,而CTGAN达11.5,且有1/9的案例出现相反推荐。一种自助基线方法在SSF上与PolicySynth相当,但直接复制真实记录并失败于成员推断防御,表明单一指标不足以评估。PolicySynth能可靠支持方向性投放筛选;其投资回报估计与真实结果偏差为70%至78%,需通过文中提出的体积校正来修正。
原文摘要 · Abstract (English)
Decision support systems (DSS) increasingly run retention what-if analysis on synthetic customer populations, because privacy constraints preclude unrestricted use of real data. Such a system is trustworthy only if the synthetic data lead managers to the same decisions as the real data would; yet prevailing criteria certify distributional similarity, not decision alignment, so a synthetic population can match every marginal distribution while still steering a marketing team toward the wrong campaigns. We close this decision-alignment gap with three contributions: strategy simulation fidelity (SSF), a criterion measuring how often the synthetic population yields the same go/no-go campaign decision as the real population; PolicySynth, a DSS framework whose generator is conditioned on the production churn scorer to align decision-relevant structure; and a three-axis reporting standard of decision alignment, membership-inference resistance, and novel-record rate as the minimum deployment quality gate. On a telecommunications churn corpus and a banking acquisition corpus, PolicySynth attains a mean SSF of 0.923 and 0.960, with seed-to-seed variance roughly ten times tighter than CTGAN on telecommunications and 2.5 times on banking. This stability is the deployable property: go/no-go recommendations shift by at most 1.2 percentage points between monthly retraining cycles, against 11.5 for CTGAN, a reversed recommendation on one campaign in nine. A bootstrap baseline matches PolicySynth on SSF yet copies real records verbatim and fails membership inference, evidence that no single axis suffices. PolicySynth reliably supports directional go/no-go screening; its ROI estimates diverge from real outcomes by 70 to 78% and require the volume correction we document.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。