arXiv:2605.14304cs.LGcs.AI2026-05

用矩阵抽象轨迹局部动态,实现高效迁移学习

Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry

  • 用正定矩阵表示轨迹片段的统计特征,捕捉局部动态结构
  • 在有限预算下达成0.73平均目标AUC,优于基线方法
  • 适合需要快速适应新任务的强化学习场景

序列决策中的组合泛化需识别先前轨迹中仍可用的部分。现有方法复用技能或预测模型,但常忽略丰富的局部转移几何与动力学。我们提出矩阵空间强化学习(MSRL),通过提升一步转移的一阶和二阶统计量构造正定矩阵描述符,表征轨迹片段。该描述符揭示共享隐结构,在抽象矩阵空间支持代数组合,并暴露迁移机会。理论证明:描述符在坐标规范下定义良好,对低阶加性信号类完全,段落组合下可加,且在所有可加描述符中最小充分。进一步表明,以轨迹片段矩阵条件化价值函数,可获得动作价值的一阶光滑近似,使源任务学习的矩阵到价值映射可引导新任务学习。MSRL可无缝集成至标准无模型与基于模型方法,阻塞过滤排除不合理组合。实验显示,MSRL在有限预算下达到0.73平均目标AUC,优于从零训练的MSRL(0.65)、TD-MPC-PT+FT(0.63)和TD-MPC(0.57)。

原文摘要 · Abstract (English)

Compositional generalization in sequential decision-making requires identifying which parts of prior rollouts remain useful for new tasks. Existing methods reuse skills or predictive models, but often overlook rich local transition geometry and dynamics. We propose Matrix-Space Reinforcement Learning (MSRL), a geometric abstraction that represents trajectory segments through positive semidefinite matrix descriptors aggregating first- and second-order statistics of lifted one-step transitions. These descriptors expose shared hidden structure, support algebraic composition in an abstract matrix space, and reveal opportunities for transfer. We prove that the descriptor is well defined up to coordinate gauge, complete for the induced low-order additive signal class, additive under valid segment composition, and minimally sufficient among admissible additive descriptors. We further show that conditioning value functions on the trajectory-segment matrix yields a first-order smooth approximation of action values, enabling source-learned matrix-to-value mappings to bootstrap learning in new tasks. MSRL is plug-in compatible with standard model-free and model-based methods, while obstruction filtering rejects implausible compositions. Empirically, MSRL achieves the best average finite-budget target AUC of 0.73, outperforming MSRL from scratch (0.65), TD-MPC-PT+FT (0.63), and TD-MPC (0.57).

强化学习迁移学习几何抽象轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。