用聚类强化学习,在社交网络中更准更快地估计干预效果。
Estimating Causal Effects in Networks with Cluster-Based Bandits
- 基于聚类的多臂赌博机,动态平衡探索与利用
- 相比传统实验,奖励收益更高,误差控制在可接受范围
- 适合存在干扰的社交网络场景,尤其关注长期优化
随机对照试验(RCT)是评估因果效应的金标准,但在社交网络等存在干扰的场景下,个体间相互影响使传统方法失效。此外,若某处理组表现差,持续分配会导致显著性能损失。为此,本文提出两种基于聚类的多臂赌博机(MAB)算法,通过逐步学习网络中的总处理效应,在最大化期望回报的同时实现自适应优化。在模拟干扰的半真实数据上对比发现:忽略聚类的朴素MAB算法虽有更高奖励-行动比,但因溢出效应导致处理效应估计误差较大;而聚类型MAB算法在保持相近估计精度的前提下,相较对应RCT方法实现了更高的奖励-行动比,表现出更强的实用性与效率。
原文摘要 · Abstract (English)
The gold standard for estimating causal effects is randomized controlled trial (RCT) or A/B testing where a random group of individuals from a population of interest are given treatment and the outcome is compared to a random group of individuals from the same population. However, A/B testing is challenging in the presence of interference, commonly occurring in social networks, where individuals can impact each others outcome. Moreover, A/B testing can incur a high performance loss when one of the treatment arms has a poor performance and the test continues to treat individuals with it. Therefore, it is important to design a strategy that can adapt over time and efficiently learn the total treatment effect in the network. We introduce two cluster-based multi-armed bandit (MAB) algorithms to gradually estimate the total treatment effect in a network while maximizing the expected reward by making a tradeoff between exploration and exploitation. We compare the performance of our MAB algorithms with a vanilla MAB algorithm that ignores clusters and the corresponding RCT methods on semi-synthetic data with simulated interference. The vanilla MAB algorithm shows higher reward-action ratio at the cost of higher treatment effect error due to undesired spillover. The cluster-based MAB algorithms show higher reward-action ratio compared to their corresponding RCT methods without sacrificing much accuracy in treatment effect estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。