GenCAR通过可控的反事实推荐提升分布外推荐可靠性
GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation

- 基于偏好锚点与可信半径筛选生成反事实候选
- 实现分布外推荐中代理标签假发现率的有限样本控制
- 适合关注推荐系统鲁棒性与风险控制的研究者
在分布外推荐场景下,平衡效用与风险至关重要。现有方法多聚焦于排序优化或生成反事实候选,但未控制代理标签假发现率(FDR)。本文提出α-有效反事实推荐(α-VCR)问题,在保留反事实监督学习支持的同时,严格控制代理标签FDR。为此,我们设计GenCAR,将基于偏好的反事实监督与校准集合选择相结合:固定稳定偏好表征,干预环境因素;通过偏好锚点与信任半径过滤离线大模型提案;使用保形p值进行Benjamini-Hochberg选择。理论上,我们界定了条件反事实近似误差,并证明在可交换性和正回归依赖条件下,可实现有限样本、无分布假设的代理标签FDR控制,且在任意依赖下具备Benjamini-Yekutieli保证。大量实验验证了真实代理假发现比例的控制效果,表明GenCAR在多种基准上持续提升分布外候选恢复能力。
原文摘要 · Abstract (English)
Serving useful recommendations under distribution shift is crucial for balancing utility and risk in out-of-distribution (OOD) recommendation. However, most existing OOD methods improve ranking or construct counterfactual candidates without controlling the proxy-label false discovery rate (FDR) of the served set. In this work, we formulate OOD serving as the $\alpha$-Valid Counterfactual Recommendation ($\alpha$-VCR) problem to retain candidate support learned from counterfactual supervision while controlling proxy-label FDR, and propose GenCAR, which couples preference-grounded counterfactual supervision with calibrated set selection. In particular, GenCAR fixes the stable-preference representation while intervening on the environmental factor, grounds offline large language model proposals through preference anchors and trust-radius filtering, and uses conformal $p$-values for Benjamini--Hochberg selection. We theoretically bound conditional counterfactual approximation error and prove finite-sample, distribution-free control of proxy-label FDR under exchangeability and positive regression dependence, with a Benjamini--Yekutieli guarantee under arbitrary dependence. Extensive experiments audit realized proxy false discovery proportions and demonstrate that GenCAR consistently enhances OOD candidate recovery across diverse benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。