首个真实工业装配多粒度数据集,支持多视角同步捕捉与异常恢复标注。
IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly

- 构建五视角同步RGB-D数据,覆盖真实装配全流程。
- 包含112次实验、39.5小时数据,支持手部动作到流程状态的分层标注。
- 适合工业视觉理解、人机协作系统研究者使用。
我们提出IMPACT,一个面向工业装配任务的同步五视角RGB-D数据集,基于真实商用角磨机的组装与拆卸过程,采用专业工具完成。据我们所知,IMPACT是首个在真实工业流程中同时提供同步自外/内视角RGB-D采集、解耦双臂标注、符合性感知状态追踪以及显式异常-恢复监督的基准数据集。数据集包含13名参与者完成的112次试验,总时长39.5小时,执行路径由部分顺序前提图控制,涵盖六类异常分类,并通过NASA-TLX测量操作员认知负荷。标注层级将手部原子动作关联至粗粒度流程步骤、组件状态及每只手的符合性阶段,各视角同步空段用于分离感知局限与算法失效。系统性基线揭示了单任务基准无法体现的根本局限,尤其在存在观测不全、路径灵活和修正行为的真实部署条件下。完整数据集、标注与评估代码已公开于https://github.com/Kratos-Wen/IMPACT。
原文摘要 · Abstract (English)
We introduce IMPACT, a synchronized five-view RGB-D dataset for deployment-oriented industrial procedural understanding, built around real assembly and disassembly of a commercial angle grinder with professional-grade tools. To our knowledge, IMPACT is the first real industrial assembly benchmark that jointly provides synchronized ego-exo RGB-D capture, decoupled bimanual annotation, compliance-aware state tracking, and explicit anomaly--recovery supervision within a single real industrial workflow. It comprises 112 trials from 13 participants totaling 39.5 hours, with multi-route execution governed by a partial-order prerequisite graph, a six-category anomaly taxonomy, and operator cognitive load measured via NASA-TLX. The annotation hierarchy links hand-specific atomic actions to coarse procedural steps, component assembly states, and per-hand compliance phases, with synchronized null spans across views to decouple perceptual limitations from algorithmic failure. Systematic baselines reveal fundamental limitations that remain invisible to single-task benchmarks, particularly under realistic deployment conditions that involve incomplete observations, flexible execution paths, and corrective behavior. The full dataset, annotations, and evaluation code are available at https://github.com/Kratos-Wen/IMPACT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。