通过离线诊断测试推荐系统对回报条件的可控性,发现局部干预效果显著。
Auditing Return Conditioning as a Control Knob: An Offline Diagnostic for Decision Transformer Recommendation
- 用不同局部性干预测试决策变换器对回报条件的响应能力
- 全上下文干预使电影类型预测提升23.61个百分点,仅改当前项仅提升1.77点
- 引入四类验证方法,为模型可控性提供可复现的审计框架
离线回报预估(RTG)扫描可用于检验基于回报的推荐系统是否可调控,但此类干预极少被审计。重写所有历史RTG令牌会生成越来越合成的上下文,而仅修改当前令牌则更局部。我们在固定窗口的离线设置中测试这一差异。在MovieLens 25M和MyAnimeList 2020(MAL)上,评估了决策变换器在RTG局部性阶梯、无RTG控制、记录匹配与评分奖励检查、以及轨迹内随机化RTG消融下的表现。在MovieLens上,仅作用于真实上下文位置的$K=20$干预,使犯罪类预测占比从验证集第5百分位升至第95百分位,提升$+23.61 \pm 2.96$个百分点;而仅改当前槽位仅提升$+1.77 \pm 1.17$点。随机化RTG模型几乎消除该响应($K=20$时仅$+2.08 \pm 1.20$点)。在MAL上,相同协议未引发戏剧类响应:$K=20$时戏剧类变化$-0.03 \pm 0.07$点,$K=1$时$-0.01 \pm 0.01$点。各类别下真实RTG、无RTG与随机化RTG的类型预测准确率相近,且$K=1$时记录匹配率与匹配评分变化极小。由于数据集与目标类别选择属探索性,这些数值仅为描述性;跨诊断模式(局部性、随机化RTG、MAL零结果)未能确立奖励控制。我们提出四项检查:干预局部性、无RTG基线、奖励检查与RTG内容消融。
原文摘要 · Abstract (English)
Offline return-to-go (RTG) sweeps can test whether a recommender conditioned on return is controllable, but the intervention is rarely audited. Rewriting every historical RTG token creates an increasingly synthetic context, while rewriting only the current token is more local. We test this distinction in an offline setting with a fixed window. On MovieLens 25M and MyAnimeList 2020 (MAL), we evaluate a Decision Transformer using an RTG locality ladder, a control without RTG, a logged match and score reward check, and a within-trajectory shuffled RTG ablation. On MovieLens, a $K=20$ intervention that covers the full context, applied only to real context positions, shifts the share of Crime predictions by $+23.61 \pm 2.96$ percentage points from the validation 5th to 95th percentile, whereas changing only the current slot shifts it by $+1.77 \pm 1.17$ points. The shuffled RTG model largely removes this response ($+2.08 \pm 1.20$ points at $K=20$). On MAL, the same protocol does not produce a Drama response: $K=20$ changes Drama by $-0.03 \pm 0.07$ points, and $K=1$ by $-0.01 \pm 0.01$. Genre prediction accuracy is numerically close across real RTG, no RTG, and shuffled RTG, and at $K=1$ logged match rates and matched ratings change little. Because dataset and focus-genre selection were exploratory, these magnitudes are descriptive; the cross-diagnostic pattern across locality, shuffled RTG, and the null result on MAL does not establish reward control. We propose four checks: intervention locality, a no-RTG baseline, a reward check, and RTG-content ablation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。