通过物理纠正实时推断人机协作中的人类任务意图
TATIC: Task-Aware Temporal Learning for Human Intent Inference from Physical Corrections in Human-Robot Collaboration
- 用扭矩估计和任务感知时序网络解析物理反馈
- 意图识别宏平均F1达0.904,硬件验证成功
- 适合需要自然人机交互的工业协作场景
在人-机器人协作(HRC)中,机器人需在线适应动态任务约束和不断变化的人类意图。尽管物理纠正提供了低延迟的运动级调整通道,但从这些短暂互动中提取任务级语义意图仍具挑战。现有基于基础模型的方法主要依赖视觉和语言输入,缺乏对物理反馈的解释能力;而传统物理人机交互(pHRI)方法虽利用物理纠正进行轨迹引导,却难以推断任务级语义。为此,我们提出TATIC,一个统一框架,结合基于扭矩的接触力估计与任务感知时序卷积网络(TCN),从简短物理纠正中联合推断离散任务意图和连续运动参数。任务对齐特征规范化确保在多种布局下的鲁棒泛化,意图驱动的自适应机制将推断出的人类意图转化为机器人动作调整。实验在意图识别上达到0.904的宏平均F1分数,并在协同拆卸任务中完成硬件验证(视频见https://youtu.be/xF8A52qwEc8)。
原文摘要 · Abstract (English)
In human-robot collaboration (HRC), robots must adapt online to dynamic task constraints and evolving human intent. While physical corrections provide a natural, low-latency channel for operators to convey motion-level adjustments, extracting task-level semantic intent from such brief interactions remains challenging. Existing foundation-model-based approaches primarily rely on vision and language inputs and lack mechanisms to interpret physical feedback. Meanwhile, traditional physical human-robot interaction (pHRI) methods leverage physical corrections for trajectory guidance but struggle to infer task-level semantics. To bridge this gap, we propose TATIC, a unified framework that utilizes torque-based contact force estimation and a task-aware Temporal Convolutional Network (TCN) to jointly infer discrete task-level intent and estimate continuous motion-level parameters from brief physical corrections. Task-aligned feature canonicalization ensures robust generalization across diverse layouts, while an intent-driven adaptation scheme translates inferred human intent into robot motion adaptations. Experiments achieve a 0.904 Macro-F1 score in intent recognition and demonstrate successful hardware validation in collaborative disassembly (see experimental video at https://youtu.be/xF8A52qwEc8).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。