对比学习批评者虽能排序动作,却无法可靠指导选择,易因嵌入幅值漂移导致错误决策。
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search
- 用余弦对比训练批评者,但嵌入幅值漂移引发非支持动作得分虚高
- 在四个OGBench导航任务中,前10%最高分项得分弱校准或反向,无法正确排序
- 适合用作动作兼容性排序,但动作选择需额外价值校准的标量信号
良好的动作排序并不保证对比批评者可安全最大化。这些批评者逐渐表现出类似价值的目标特性,用于最佳K选、规划及批评引导生成。无界双线性得分会导致大嵌入范数放大非支持动作的值,而余弦归一化无法消除此缺陷。受控支持分解表明,大部分原始双线性遗憾主要源于范数漂移。尽管如此,余弦与混合批评者仍从多数池中选出非支持动作,并产生相近遗憾。在四个OGBench导航任务中,对比得分在前十分位表现弱校准或反转,且无法按价值对固定查询动作进行排序。贝尔曼训练的TD-Q表现成功,包括在参数匹配的函数类对照实验中。实际代价依赖任务:模拟回放显示在PointMaze上单步选择代价显著,而在AntMaze和HumanoidMaze上则为充分幂的零假设,因控制器具备自我纠正能力。训练/读出分解揭示了排序丢失源于余弦训练目标;未经训练的嵌入在推理时归一化后仍保留弱排序能力。因此,候选最大化可能利用由范数漂移、得分饱和或支持内错序造成的假阳性。对比批评者在导航与操作任务中仍可作为有用的兼容性排序器,但动作选择需依赖价值校准的标量信号。
原文摘要 · Abstract (English)
Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-$K$ selection, planning, and critic-guided generation. Unbounded bilinear scores can let large embedding norms inflate off-support values, but cosine bounding does not remove the failure. A controlled support decomposition attributes most raw bilinear regret to norm drift. Cosine and hybrid critics nevertheless select off-support actions from most pools and incur comparable regret. Contrastive scores are weakly calibrated or inverted in the top score decile across four OGBench navigation tasks, and they fail to order fixed-query actions by value. Bellman-trained TD-Q succeeds, including in a parameter-matched function-class control. Realized costs depend on the task: simulator rollouts reveal single-step selection costs on PointMaze and the exact-$Q^*$ toy but well-powered nulls on AntMaze and HumanoidMaze, where the controller can self-correct. A training/readout decomposition traces the lost ordering to the cosine training objective; raw-trained embeddings retain weak ordering after inference-time normalization. Candidate maximization can therefore exploit false positives caused by norm drift, score saturation, or in-support misranking. Contrastive critics remain useful compatibility rankers on navigation and manipulation tasks, but action selection requires a value-calibrated scalar.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。