arXiv:2608.21778cs.RO2026-08中稿 · ISCSIC 2026

用视觉目标控制挖机,让机器更懂人想挖哪里。

Vision Guided Target Conditioned Control for Autonomous Excavation

论文配图:Vision Guided Target Conditioned Control for Autonomous Excavation
图 1 · 摘自论文原文
  • 用图像掩码告诉机器要挖哪块地,结合多视角图像和自身状态生成连续操作指令。
  • 在挖土任务中,正确率从15.7%提升至96%,清土量达76.8%,接近人类效率。
  • 适合想做智能工程机械自动化、特别是基于视觉控制的研究者。

自主挖掘需要智能控制系统,将空间作业意图转化为在高接触力土壤交互下的协调铲斗运动。本文提出一种基于物理模拟的可变形土壤环境中的目标条件化智能控制框架。以对齐图像的目标掩码作为视觉空间指令,指定期望挖掘区域;采用掩码条件化的动作分块变换器(ACT),将多视角RGB图像、本体感知与目标掩码映射为时序扩展的操纵杆指令。为减少忽略目标的行为,演示数据采用成对条件监督:同一场景下使用不同目标掩码及对应动作块进行示范。框架通过诊断性操作任务与单铲及连续堆土清理协议的挖掘仿真基准进行评估。在操作任务中,无条件ACT成功率为4%,非配对掩码条件化ACT为63%,配对条件化ACT达96%。在连续堆土清理中,配对条件化ACT清除76.8%的土堆,远超两个基线(27.4%与15.7%),人类效率归一化达91.0%。结果表明,视觉目标条件化、成对示范结构与动作分块控制共同构成了一条实用的虚实结合挖机自动化流水线。

原文摘要 · Abstract (English)

Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4\% for no-condition ACT, 63\% for non-paired mask-conditioned ACT, and 96\% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8\% of the pile versus 27.4\% and 15.7\% for the two baselines, with 91.0\% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.

自主挖掘视觉控制强化学习机器人仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。