让机器同时理解与执行动作,提升智能体的感知与行为协同能力。
Embodied Representation Alignment with Mirror Neurons
- 通过共享隐空间和对比学习,显式对齐观察与执行动作的中间表示。
- 在多个数据集上实现动作识别与执行任务的双向性能提升。
- 适合研究具身智能、动作理解与神经机制启发的模型设计者。
镜像神经元是一类在个体观察动作和执行相同动作时均会激活的神经元,揭示了动作理解与具身执行之间存在根本性关联。然而,现有机器学习方法大多忽略这一关联,将两者视为独立任务。本文从表征学习视角提出统一建模框架,首次发现动作观察与执行的中间表示会自发对齐。受镜像神经元启发,进一步引入显式对齐方法:使用两个线性层将两类表示映射至共享隐空间,通过对比学习强制对应表示对齐,有效最大化其互信息。实验表明,该简单方法在多个数据集上实现了两任务间的相互促进,显著提升表征质量与泛化能力。
原文摘要 · Abstract (English)
Mirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodied execution, suggesting that these two abilities are inherently connected. Nonetheless, existing machine learning methods largely overlook this interplay, treating these abilities as separate tasks. In this study, we provide a unified perspective in modeling them through the lens of representation learning. We first observe that their intermediate representations spontaneously align. Inspired by mirror neurons, we further introduce an approach that explicitly aligns the representations of observed and executed actions. Specifically, we employ two linear layers to map the representations to a shared latent space, where contrastive learning enforces the alignment of corresponding representations, effectively maximizing their mutual information. Experiments demonstrate that this simple approach fosters mutual synergy between the two tasks, effectively improving representation quality and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。