收集七小时真人操作视频,帮助机器人理解家居任务中的动作可能性。
The Wilhelm Tell Dataset of Affordance Demonstrations
- 用第一/第三人称视频记录日常任务操作过程
- 包含任务中展现的动作可能性元数据,总时长约7小时
- 适合研究人机协作与任务准备行为的机器人开发
动作可能性(affordances)——即环境或物体提供的可操作性——对在人类环境中运行的机器人感知至关重要。现有方法通常基于标注的静态图像或形状训练这些能力。本文提出一个新型数据集,用于常见家庭任务的动作可能性学习。与以往方法不同,本数据集包含从第一人称和第三人称视角记录的任务视频序列,附带任务中体现的动作可能性元数据,旨在训练感知系统识别动作可能性的表现。数据由多名参与者采集,总计约七小时的人类活动记录。任务表现的多样性还支持研究人们为任务所做的准备动作,如如何布置操作空间,这对协作服务机器人具有重要意义。
原文摘要 · Abstract (English)
Affordances - i.e. possibilities for action that an environment or objects in it provide - are important for robots operating in human environments to perceive. Existing approaches train such capabilities on annotated static images or shapes. This work presents a novel dataset for affordance learning of common household tasks. Unlike previous approaches, our dataset consists of video sequences demonstrating the tasks from first- and third-person perspectives, along with metadata about the affordances that are manifested in the task, and is aimed towards training perception systems to recognize affordance manifestations. The demonstrations were collected from several participants and in total record about seven hours of human activity. The variety of task performances also allows studying preparatory maneuvers that people may perform for a task, such as how they arrange their task space, which is also relevant for collaborative service robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。