arXiv:2410.02921cs.CV2024-10被引 2

人类在空中画字母的视频数据集,挑战视觉模型对复杂动作的识别能力。

AirLetters: An Open Video Dataset of Characters Drawn in the Air

  • 采集真实人类空中画字母的视频,强调运动模式与长期时序信息融合。
  • 现有模型表现远低于人类基准,识别准确率显著落后。
  • 适合研究动作识别、时序建模和具身智能的学者参考。

我们提出 AirLetters,一个包含人类在空中绘制字母的真实世界视频数据集。该任务要求视觉模型预测人类在空中绘制的字母,其准确分类依赖于对运动模式的辨别以及视频中长时序信息的整合。对当前主流图像与视频理解模型在 AirLetters 上的广泛评估表明,这些方法表现不佳,远低于人类基准。我们的工作揭示:尽管端到端视频理解取得进展,但对复杂人体动作的准确建模——对人类而言是直觉任务——仍是端到端学习中的开放难题。

原文摘要 · Abstract (English)

We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict letters that humans draw in the air. Unlike existing video datasets, accurate classification predictions for AirLetters rely critically on discerning motion patterns and on integrating long-range information in the video over time. An extensive evaluation of state-of-the-art image and video understanding models on AirLetters shows that these methods perform poorly and fall far behind a human baseline. Our work shows that, despite recent progress in end-to-end video understanding, accurate representations of complex articulated motions -- a task that is trivial for humans -- remains an open problem for end-to-end learning.

动作识别视频理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。