arXiv:2604.26637cs.ROcs.AI2026-04被引 1

ATLAS是一款专为机器人长时序动作分割设计的标注工具,支持多模态数据同步标注。

ATLAS: An Annotation Tool for Long-horizon Robotic Action Segmentation

论文配图:ATLAS: An Annotation Tool for Long-horizon Robotic Action Segmentation
图 1 · 摘自论文原文
  • 支持视频与机器人本体信号(如夹爪状态)的时序同步可视化
  • 在装配任务中将单动作标注时间减少6%以上,边界误差降低五倍
  • 适用于多种机器人数据格式,适合研究者高效构建高质量标注数据

对长时序机器人演示进行精确的时间动作边界标注,对于训练和评估动作分割与操作策略学习方法至关重要。现有标注工具通常存在局限:主要针对视觉数据设计,不原生支持机器人特有的时间序列信号(如夹爪状态或力/扭矩)同步可视化,且需大量工作适配不同数据集格式。本文提出ATLAS,一款专为长时序机器人动作分割设计的标注工具。ATLAS提供多模态机器人数据(包括多视角视频与本体感知信号)的时序同步可视化,并支持动作边界、动作标签及任务结果的标注。工具原生支持广泛使用的机器人数据格式,如ROS bag和强化学习数据集(RLDS)格式,并直接支持REASSEMBLE等特定数据集。通过模块化数据抽象层,可轻松扩展至新格式。其以键盘为中心的界面最大限度减少标注负担,提升效率。在接触丰富的装配任务实验中,相比ELAN工具,平均每个动作标注时间至少减少6%;引入时间序列数据后,与专家标注的时序对齐度提升超过2.8%,边界误差相比仅依赖视觉的工具下降五倍。

原文摘要 · Abstract (English)

Annotating long-horizon robotic demonstrations with precise temporal action boundaries is crucial for training and evaluating action segmentation and manipulation policy learning methods. Existing annotation tools, however, are often limited: they are designed primarily for vision-only data, do not natively support synchronized visualization of robot-specific time-series signals (e.g., gripper state or force/torque), or require substantial effort to adapt to different dataset formats. In this paper, we introduce ATLAS, an annotation tool tailored for long-horizon robotic action segmentation. ATLAS provides time-synchronized visualization of multi-modal robotic data, including multi-view video and proprioceptive signals, and supports annotation of action boundaries, action labels, and task outcomes. The tool natively handles widely used robotics dataset formats such as ROS bags and the Reinforcement Learning Dataset (RLDS) format, and provides direct support for specific datasets such as REASSEMBLE. ATLAS can be easily extended to new formats via a modular dataset abstraction layer. Its keyboard-centric interface minimizes annotation effort and improves efficiency. In experiments on a contact-rich assembly task, ATLAS reduced the average per-action annotation time by at least 6% compared to ELAN, while the inclusion of time-series data improved temporal alignment with expert annotations by more than 2.8% and decreased boundary error fivefold compared to vision-only annotation tools.

机器人标注动作分割多模态数据数据工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。