arXiv:2510.11321cs.RO2025-10NeurIPS被引 2

无需标注数据,自动挖掘机器人操作的层次化概念

HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data

  • 通过跨模态相关性与多时序抽象,自监督学习操作概念
  • 在仿真和真实场景中提升策略性能,概念具可解释性
  • 适合研究机器人表示学习与强化学习的开发者

机器人操作的有效泛化依赖于捕捉跨环境与任务的不变交互模式。本文提出一种自监督框架,通过跨模态感官相关性与多时序层次抽象,从无标注多模态数据中学习层次化操作概念,无需人工标注。方法结合跨模态相关网络识别感官间持续模式,以及多时程预测器构建不同时间尺度的层次化表征。所学概念使策略聚焦可迁移的关系模式,同时兼顾即时动作与长期目标。在多个仿真基准和真实部署中验证,概念增强策略显著提升性能。分析显示,学习到的概念虽无语义监督,但与人类可理解的操作原语相似。本工作推进了操作表示学习的理解,并提供复杂场景下提升机器人性能的实用方案。

原文摘要 · Abstract (English)

Effective generalization in robotic manipulation requires representations that capture invariant patterns of interaction across environments and tasks. We present a self-supervised framework for learning hierarchical manipulation concepts that encode these invariant patterns through cross-modal sensory correlations and multi-level temporal abstractions without requiring human annotation. Our approach combines a cross-modal correlation network that identifies persistent patterns across sensory modalities with a multi-horizon predictor that organizes representations hierarchically across temporal scales. Manipulation concepts learned through this dual structure enable policies to focus on transferable relational patterns while maintaining awareness of both immediate actions and longer-term goals. Empirical evaluation across simulated benchmarks and real-world deployments demonstrates significant performance improvements with our concept-enhanced policies. Analysis reveals that the learned concepts resemble human-interpretable manipulation primitives despite receiving no semantic supervision. This work advances both the understanding of representation learning for manipulation and provides a practical approach to enhancing robotic performance in complex scenarios.

机器人自监督层次表征多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。