辅助动态监督无法提升人形机器人模拟中的表征可解释性与鲁棒性
Evaluating Factor-Wise Auxiliary Dynamics Supervision for Latent Structure and Robustness in Simulated Humanoid Locomotion
- 通过分因子辅助损失训练24维潜在变量,试图构建可解码的结构化表示
- 探针测试显示潜在空间无功能分离特征(R²≈0),且对奖励无显著提升
- 循环网络虽表现更好,但优势源于瓶颈压缩而非辅助监督
我们评估了分因子辅助动态监督是否能生成有用的潜在结构或提升模拟人形机器人行走的鲁棒性。DynaMITE——一个使用因子分解24维潜在变量并通过近端策略优化(PPO)中各因子辅助损失训练的Transformer编码器——在Isaac Lab的四个任务上与LSTM、普通Transformer和MLP基线进行对比。结果表明,受监督的潜在空间未表现出可解码或功能分离的因子结构:所有五个动力学因子的探针R²均接近0,子空间变化对奖励影响小于0.05,标准解耦指标(MIG、DCI、SAP)也接近零。无监督LSTM隐藏状态的探针R²更高(最高达0.10)。2×2因子消融实验(n=10种子)显示,辅助损失对分布内(ID)奖励(+0.03,p=0.732)和严重分布外(OOD)奖励(+0.03,p=0.669)无显著影响,而tanh瓶颈在两种情形下均有小幅稳定优势(ID:+0.16,p=0.207;OOD:+0.10,p=0.208)。该优势在严重联合扰动下仍存在但不增强,说明是训练期表示优势,而非鲁棒性机制。LSTM在所有四任务上的名义奖励最优(p<0.03);DynaMITE在联合偏移压力下退化更少(2.3% vs. 16.7%),但此差异归因于瓶颈压缩,非辅助监督。对运动控制研究者而言:辅助动态监督无法生成可解释估计量,也未显著提升奖励或鲁棒性,仅优于瓶颈本身;循环基线仍是名义性能更强的选择。
原文摘要 · Abstract (English)
We evaluate whether factor-wise auxiliary dynamics supervision produces useful latent structure or improved robustness in simulated humanoid locomotion. DynaMITE -- a transformer encoder with a factored 24-d latent trained by per-factor auxiliary losses during proximal policy optimization (PPO) -- is compared against Long Short-Term Memory (LSTM), plain Transformer, and Multilayer Perceptron (MLP) baselines on a Unitree G1 humanoid across four Isaac Lab tasks. The supervised latent shows no evidence of decodable or functionally separable factor structure: probe R^2 ~ 0 for all five dynamics factors, clamping any subspace changes reward by < 0.05, and standard disentanglement metrics (MIG, DCI, SAP) are near zero. An unsupervised LSTM hidden state achieves higher probe R^2 (up to 0.10). A 2x2 factorial ablation (n = 10 seeds) isolates the contributions of the tanh bottleneck and auxiliary losses: the auxiliary losses show no measurable effect on either in-distribution (ID) reward (+0.03, p = 0.732) or severe out-of-distribution (OOD) reward (+0.03, p = 0.669), while the bottleneck shows a small, consistent advantage in both regimes (ID: +0.16, p = 0.207; OOD: +0.10, p = 0.208). The bottleneck advantage persists under severe combined perturbation but does not amplify, indicating a training-time representation benefit rather than a robustness mechanism. LSTM achieves the best nominal reward on all four tasks (p < 0.03); DynaMITE degrades less under combined-shift stress (2.3% vs. 16.7%), but this difference is attributable to the bottleneck compression, not the auxiliary supervision. For locomotion practitioners: auxiliary dynamics supervision does not produce an interpretable estimator and does not measurably improve reward or robustness beyond what the bottleneck alone provides; recurrent baselines remain the stronger choice for nominal performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。