arXiv:2508.07842cs.ROcs.AI2025-08被引 7

提出跨域长时任务学习框架,解耦环境与技能信息以提升泛化能力。

DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts

  • 双流解耦设计:分离环境感知与自身状态处理
  • 跨域任务成功率平均提升23%,执行效率提高29%
  • 适合需要灵活组合技能的复杂交互场景研究

人类-场景交互中的长时程任务是复杂的多步任务,需持续规划、顺序决策并跨域执行。现有方法依赖预训练子任务拼接,环境观测与自身状态紧密耦合,难以泛化到新环境与技能组合,无法完成跨域长时任务。为此,本文提出DETACH框架,受大脑“何处-何物”双通路启发,采用双模块设计:1)环境学习模块,捕捉物体功能、空间关系与场景语义,通过完全解耦实现跨域迁移;2)技能学习模块,处理关节自由度与运动模式等自身状态,独立编码运动模式以实现跨技能迁移。在多种长时任务上进行实验,相比现有方法,平均子任务成功率提升23%,执行效率提高29%。

原文摘要 · Abstract (English)

Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across domains to achieve the final goal. However, existing methods heavily rely on skill chaining by concatenating pre-trained subtasks, with environment observations and self-state tightly coupled, lacking the ability to generalize to new combinations of environments and skills, failing to complete various LH tasks across domains. To solve this problem, this paper presents DETACH, a cross-domain learning framework for LH tasks via biologically inspired dual-stream disentanglement. Inspired by the brain's "where-what" dual pathway mechanism, DETACH comprises two core modules: i) an environment learning module for spatial understanding, which captures object functions, spatial relationships, and scene semantics, achieving cross-domain transfer through complete environment-self disentanglement; ii) a skill learning module for task execution, which processes self-state information including joint degrees of freedom and motor patterns, enabling cross-skill transfer through independent motor pattern encoding. We conducted extensive experiments on various LH tasks in HSI scenes. Compared with existing methods, DETACH can achieve an average subtasks success rate improvement of 23% and average execution efficiency improvement of 29%.

长时任务跨域迁移解耦学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。