提出ALAS框架,让机器人跨场景更智能地完成复杂长程任务。
ALAS: Adaptive Long-Horizon Action Synthesis via Async-pathway Stream Disentanglement

- 模仿大脑双通路机制,分离环境信息与自身状态
- 在多个场景中提升任务成功率23%,执行效率提高29%
- 适合需要跨场景、跨技能泛化的机器人任务研究
人类-场景交互中的长时序任务是复杂的多步任务,需持续规划、顺序决策并长时间执行才能达成目标。现有方法依赖预训练子任务拼接,环境观察与自我状态紧密耦合,难以泛化到新场景和新技能组合,无法完成跨领域多样化长时序任务。为此,本文提出ALAS框架,基于生物启发的双通路解耦机制,实现跨领域学习。该框架包含两个核心模块:一是环境学习模块,用于空间理解,捕捉物体功能、空间关系与场景语义,通过完全的环境-自我解耦实现跨域迁移;二是技能学习模块,处理包括关节自由度与运动模式在内的自我状态信息,通过独立运动模式编码实现跨技能迁移。我们在多种人-场景交互长时序任务上进行了广泛实验。相比现有方法,ALAS平均子任务成功率提升23%,平均执行效率提升29%。
原文摘要 · Abstract (English)
Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across domains to achieve the final goal. However, existing methods heavily rely on skill chaining by concatenating pre-trained subtasks, with environment observations and self-state tightly coupled, lacking the ability to generalize to new combinations of environments and skills, failing to complete various LH tasks across domains. To solve this problem, this paper presents ALAS, a cross-domain learning framework for LH tasks via biologically inspired dual-stream disentanglement. Inspired by the brain's "where-what" dual pathway mechanism, ALAS comprises two core modules: i) an environment learning module for spatial understanding, which captures object functions, spatial relationships, and scene semantics, achieving cross-domain transfer through complete environment-self disentanglement; ii) a skill learning module for task execution, which processes self-state information including joint degrees of freedom and motor patterns, enabling cross-skill transfer through independent motor pattern encoding. We conducted extensive experiments on various LH tasks in HSI scenes. Compared with existing methods, ALAS can achieve an average subtasks success rate improvement of 23\% and average execution efficiency improvement of 29\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。