评估机器人控制适配器是否值得加,看真实收益能否从部署数据中恢复。
When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

- 通过四类收益分解判断适配器必要性:全局性能、同状态潜力、部署增益、状态分配增益。
- 实测显示仅1.34%部署增益,且0.09%分配增益下限,多数场景不达标。
- 适合对机器人策略优化有严谨验证需求的研究者和工程师。
为判断在冻结的命令控制型运动策略上添加可学习适配器是否值得,需确保接口暴露的改进既真实又可从部署时观测中恢复。本文提出一种适配器必要性审计方法,将全局运行点增益、同状态反事实潜力、相对于交叉拟合固定动作的部署增益、以及相对于频率匹配随机策略的状态分配增益进行分离。源簇学习器重新拟合这些量并约束违反情况,做出“通过/不通过/暂不确定”的决策。闭环命令-响应识别提供可选决策特征。在Go2机器人上,历史尺度前缀诊断发现5.2%同状态潜力,但仅有0.55%可恢复的分配增益。确认性审计对三种查询分布(直接控制、VGCC、MPC)在二十个独立聚类上各执行200次完整学习重拟合,设定1%部署与分配阈值、5%违规容忍度后,直接控制返回不通过,而VGCC与MPC返回暂不确定。其中VGCC平均部署增益达1.34%,但其分配下界仅为0.09%,违规上界高达6.25%。代表部署的二十簇H1审计同样返回不通过,而学习器级合成控制返回通过。该审计测试可观测信号是否足以支持状态依赖适配,而非预设适配器有价值。
原文摘要 · Abstract (English)
Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state counterfactual headroom, deployment gain over a cross-fitted fixed action, and state-allocation gain over a frequency-matched randomized policy. Source-cluster learner refits map these quantities and constraint violations to a GO/NO-GO/ABSTAIN decision. Closed-loop command- response identification provides optional decision features. On Go2, an archived scale-prefix diagnostic finds 5.2% same-state headroom but only 0.55% recovered allocation gain. Our confirmatory audit evaluates direct, scale, heading, and yaw interventions on twenty independent clusters for each of three query distributions induced by direct control, VGCC, and MPC, using 200 full learner refits. At 1% deployment and allocation thresholds and a 5% violation tolerance, direct queries return NO-GO, while VGCC and MPC queries ABSTAIN. VGCC has the largest mean deployment gain (1.34%), but its allocation lower bound is 0.09% and its violation upper bound is 6.25%. A deployment-representative twenty-cluster H1 audit also returns NO-GO, whereas a learner-level synthetic control returns GO. The audit therefore tests whether observable signal justifies state-dependent adaptation rather than presuming that an adapter is valuable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。