arXiv:2511.18127cs.CV2025-11被引 1

实时流式预测手部动作,支持语言指令引导。

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

  • 基于流式自回归架构,融合语言指令实时预测手部状态。
  • 在3D手部预测任务上提升35.8%,下游操作任务成功率提高13.4%。
  • 首个含语言指令的3D手部视频数据集,适合机器人与AR应用。

实时3D手部预测是增强现实和辅助机器人等场景中流畅人机交互的关键。然而,现有方法通常依赖离线视频序列,无法融入传达任务意图的语言指导。为此,我们提出SFHand,首个面向语言引导的3D手部流式预测框架。SFHand从连续视频流与语言指令中,自回归地预测包含手型、2D边界框、3D姿态及轨迹在内的完整未来手部状态。该框架结合流式自回归结构与区域增强记忆层,捕捉时序上下文并聚焦关键手部区域。为推动研究,我们还构建了EgoHaFL,首个同步标注3D手部姿态与语言指令的大规模数据集。实验表明,SFHand在3D手部预测上达到新SOTA,性能领先达35.8%;其学习表征可有效迁移至下游具身操作任务,在多个基准上提升任务成功率最高达13.4%。数据集:https://huggingface.co/datasets/ut-vision/EgoHaFL,项目页:https://github.com/ut-vision/SFHand。

原文摘要 · Abstract (English)

Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However, existing methods are ill-suited for these scenarios, as they typically require offline access to accumulated video sequences and cannot incorporate language guidance that conveys task intent. To overcome these limitations, we introduce SFHand, the first streaming framework for language-guided 3D hand forecasting. SFHand autoregressively predicts a comprehensive set of future 3D hand states, including hand type, 2D bounding box, 3D pose, and trajectory, from a continuous stream of video and language instructions. Our framework combines a streaming autoregressive architecture with an ROI-enhanced memory layer, capturing temporal context while focusing on salient hand-centric regions. To enable this research, we also introduce EgoHaFL, the first large-scale dataset featuring synchronized 3D hand poses and language instructions. We demonstrate that SFHand achieves new state-of-the-art results in 3D hand forecasting, outperforming prior work by a significant margin of up to 35.8%. Furthermore, we show the practical utility of our learned representations by transferring them to downstream embodied manipulation tasks, improving task success rates by up to 13.4% on multiple benchmarks. Dataset page: https://huggingface.co/datasets/ut-vision/EgoHaFL, project page: https://github.com/ut-vision/SFHand.

3D手部预测具身智能语言引导流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。