arXiv:2603.09532stat.MLcs.LG2026-03

提出BRACE算法,解决非合规推荐中推荐与实际治疗的差异问题。

What Do We Care About in Bandits with Noncompliance? BRACE: Bandits with Recommendations, Abstention, and Certified Effects

  • 设计无参数分阶段加倍算法,仅在矩阵认证后进行工具变量反演
  • 实现推荐与治疗策略的固定间隙识别,且在同质条件下可同时保证有效性
  • 适用于有私有信息、弱识别或额外工具变量等复杂场景

非合规情境下,推荐与实际执行的处理行为分离,学习目标需明确选择。平台可能关注当前中介流程中的推荐福利、未来直接控制下的治疗学习,或任意时间有效的不确定性。这些目标未必一致。本文形式化目标选择问题,识别推荐与治疗目标重合的直接控制情形,并通过实例证明:当下游参与者使用私有信息时,推荐福利可严格优于所有可测量的治疗策略。针对有限上下文平方-工具变量(square-IV)问题,提出无需调参的BRACE算法——仅在矩阵认证后执行工具变量反演,其余时间返回全范围但诚实的结构区间。该算法在上下文同质与可逆条件下,实现政策值的同步有效性、操作最优推荐策略的固定间隙识别,以及结构最优治疗策略的固定间隙识别。实验覆盖直接控制、中介现与未来的权衡、弱识别、同质性失效及矩形过度识别等情形,验证了安全性在简单问题表现为遗憾,在弱识别下体现为弃权与宽区间,在同质性失效下促使偏好推荐福利,而在额外工具变量可用时带来更紧的结构不确定性。对丰富上下文,推导出正交评分,其条件偏差可分解为依从性模型与结果模型误差之积,揭示了任意时间有效半参数工具变量推断所需稳定的成分。

原文摘要 · Abstract (English)

Bandits with noncompliance separate the learner's recommendation from the treatment actually delivered, so the learning target itself must be chosen. A platform may care about recommendation welfare in the current mediated workflow, treatment learning for a future direct-control regime, or anytime-valid uncertainty for one of those targets. These objectives need not agree. We formalize this objective-choice problem, identify the direct-control regime in which recommendation and treatment objectives collapse, and show by example that recommendation welfare can strictly exceed every learner-measurable treatment policy when downstream actors use private information. For finite-context square-IV problems we propose BRACE, a parameter-free phase-doubling algorithm that performs IV inversion only after matrix certification and otherwise returns full-range but honest structural intervals. BRACE delivers simultaneous policy-value validity, fixed-gap identification of the operationally optimal recommendation policy, and fixed-gap identification of the structurally optimal treatment policy under contextual homogeneity and invertibility. We complement the theory with a finite-context empirical benchmark spanning direct control, mediated present-versus-future tradeoffs, weak identification, homogeneity failure, and rectangular overidentification. The experiments show that safety appears as regret on easy problems, as abstention and wide valid intervals under weak identification, as a reason to prefer recommendation welfare under homogeneity failure, and as tighter structural uncertainty when extra instruments are available. For rich contexts, we also derive an orthogonal score whose conditional bias factorizes into compliance-model and outcome-model errors, clarifying what must be stabilized for anytime-valid semiparametric IV inference.

强化学习因果推断工具变量非合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。