构建了人与人协作装配齿轮箱的多模态数据集,精准记录求助、干预与结果的对应关系。
DYAD: A Multimodal Dataset of Co-Located Human Assistance

- 同步采集语音与动作,关联求助请求与任务状态变化
- 覆盖528个任务阶段和611次求助,共851条有效干预记录
- 适合研究具身助手的决策与响应机制,尤其关注协同行为建模
一个在旁协助的具身助手需跟踪任务状态、识别求助信号、选择干预方式并生成适当回应。现有程序类数据集详述个体执行过程,交互类数据集则捕捉远程语音指令或模糊协作,但缺乏将共处协助者的语言与物理干预,与执行者请求、任务状态、援助触发及结果进行联合标注。本文提出DYAD(双人协助数据集),为齿轮箱装配过程中人与人协作的同步多模态记录。在20次会话中,一名训练助手遵循“先指导后协助”策略,协助佩戴HoloLens 2的用户。DYAD将528个任务阶段与611次执行者请求,关联至851条有效的言语与物理援助记录。标注涵盖援助全流程;三个基准任务评估特定组件而非端到端系统:因果步骤理解、前触发模式预测、讲师响应生成。在829个符合条件的模式事件上,最强四种子RGB平均宏F1为0.548 ± 0.007;因果元数据达0.624,特权触发映射达0.915,表明预触发RGB无法恢复的信息。本工作贡献不在于规模,而在于建立从求助到干预、执行与结果的链式交互结构,融合第一视角与工作区感知。
原文摘要 · Abstract (English)
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They do not jointly link a co-located helper's verbal and physical interventions to performer requests, task state, assistance triggers, and outcomes. We introduce DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly. Across 20 sessions, one trained helper follows a guidance-first policy while assisting HoloLens 2 wearers. DYAD links 528 task-step intervals and 611 performer requests with 851 valid assistance records spanning verbal and physical help. DYAD's annotations span the assistance process; three reference tasks evaluate selected components rather than an end-to-end system: causal step understanding, pre-onset mode anticipation, and instructor response generation. On 829 eligible mode events, the strongest four-seed RGB mean is 0.548 +/- 0.007 macro-F1; causal metadata reaches 0.624 and a privileged trigger mapping 0.915, revealing information not recovered from pre-onset RGB. DYAD's contribution is not scale, but a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。