arXiv:2508.18802cs.RO2025-08被引 3

让机器人视觉随任务动态调整,更像人一样专注关键信息。

HyperTASR: Hypernetwork-Driven Task-Aware Scene Representations for Robust Manipulation

  • 用超网络根据任务目标和执行阶段动态生成视觉特征变换参数
  • 在仿真与真实环境均显著提升抓取成功率,优于传统静态表示方法
  • 适合需要灵活感知的机器人操控场景,尤其对复杂任务效果突出

机器人操控的有效策略学习依赖于能选择性捕捉任务相关环境特征的场景表征。现有方法通常采用任务无关的表征提取,无法模拟人类认知中的动态感知适应。我们提出HyperTASR,一种由超网络驱动的框架,根据任务目标和执行阶段动态调节场景表征。其架构会基于任务说明和进展状态动态生成表征变换参数,使表征在任务执行过程中持续演化。该方法在保持与现有策略学习框架兼容的同时,从根本上重构了视觉特征的处理方式。不同于简单拼接或融合任务嵌入与任务无关表征的方法,HyperTASR在计算上分离了任务上下文与状态依赖的处理路径,提升了学习效率与表征质量。在仿真与真实环境中的全面评估显示,该方法在多种表征范式下均有显著性能提升。通过消融实验与注意力可视化,验证了该方法能有选择地优先关注任务相关的场景信息,接近人类在操控任务中的自适应感知。项目网站:https://lisunphil.github.io/HyperTASR_projectpage/

原文摘要 · Abstract (English)

Effective policy learning for robotic manipulation requires scene representations that selectively capture task-relevant environmental features. Current approaches typically employ task-agnostic representation extraction, failing to emulate the dynamic perceptual adaptation observed in human cognition. We present HyperTASR, a hypernetwork-driven framework that modulates scene representations based on both task objectives and the execution phase. Our architecture dynamically generates representation transformation parameters conditioned on task specifications and progression state, enabling representations to evolve contextually throughout task execution. This approach maintains architectural compatibility with existing policy learning frameworks while fundamentally reconfiguring how visual features are processed. Unlike methods that simply concatenate or fuse task embeddings with task-agnostic representations, HyperTASR establishes computational separation between task-contextual and state-dependent processing paths, enhancing learning efficiency and representational quality. Comprehensive evaluations in both simulation and real-world environments demonstrate substantial performance improvements across different representation paradigms. Through ablation studies and attention visualization, we confirm that our approach selectively prioritizes task-relevant scene information, closely mirroring human adaptive perception during manipulation tasks. The project website is at https://lisunphil.github.io/HyperTASR_projectpage/.

机器人操控动态感知超网络场景表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。