arXiv:2603.07484cs.RO2026-03

分层视觉清理框架提升复杂杂乱环境下的双臂操作成功率

HSC-VLA: Hierarchical Scene-Clearing for Robust Bimanual Manipulation in Dense Clutter

  • 用高层视觉语义分解任务并生成任务相关掩码,过滤无关干扰
  • 在高密度杂乱场景中整体成功率达86.7%,较单体模型提升52.4%
  • 适合复杂长程操作任务,尤其擅长杂乱物品分类与补货场景

现代视觉-语言-动作模型在高密度操作环境中常因任务无关视觉干扰导致指令执行失败,注意力被稀释、语义定位失效,严重降低复杂长程任务表现。为突破单体端到端架构的表征瓶颈,本文提出HSC-VLA,一种分层框架,通过显式场景清理抽象将高层视觉语义推理与低层高频传感运动执行解耦。高层Brain负责分解长程任务并生成任务特定场景掩码,保留任务相关几何信息同时抑制干扰物;滤波后的观测输入底层Cerebellum——基于扩散模型的双臂操作策略,仅使用掩码过滤后的视觉与本体感觉进行控制。在密集杂乱超市货架上的大量实验表明,HSC-VLA在高密度杂乱环境下整体成功率达到86.7%,超越最优单体基线(π₀-Full FT)的34.3%,提升52.4%。其在长程任务中表现优异,杂乱物品分类成功率达72%,补货任务达66%,展现出强鲁棒性与有效故障恢复能力。

原文摘要 · Abstract (English)

Modern Vision--Language--Action models often suffer from critical instruction-following failures in high-density manipulation environments, where task-irrelevant visual clutter dilutes attention, corrupts grounding, and substantially degrades performance in complex long-horizon scenarios. To overcome the representation bottleneck of monolithic end-to-end architectures, we propose HSC-VLA, a hierarchical framework that decouples high-level visual-semantic reasoning from low-level, high-frequency sensorimotor execution through an explicit scene-clearing abstraction. HSC-VLA employs a high-level Brain to decompose long-horizon tasks and to generate task-specific scene masks that preserve task-relevant geometry while suppressing distractors. The filtered observations are then passed to a low-level Cerebellum, a diffusion-based policy that performs bimanual manipulation using only mask-filtered vision and proprioception. Extensive experiments in densely cluttered supermarket shelves demonstrate that HSC-VLA achieves 86.7\% aggregate success under high-density clutter, surpassing the best monolithic baseline ($π_0$-Full FT at 34.3\%) by 52.4\%. HSC-VLA also exhibits strong long-horizon performance, reaching 72\% on clutter sorting and 66\% on restocking, demonstrating strong robustness and effective failure recovery in complex cluttered manipulation.

双臂操作视觉清理长程任务扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。