首个微重力下人类行为与场景理解数据集,助力太空视觉系统发展
Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
- 构建真实太空任务与影视模拟结合的微重力视频数据集
- 包含4759段视频、50类动作、7000+问答对,支持多任务评估
- 适合航天视觉、人机交互及跨域鲁棒性研究者使用
尽管视频理解取得显著进展,但现有数据集大多局限于地球重力环境。而微重力会改变人体运动、交互方式和视觉语义,暴露出现实视觉系统在安全关键型太空应用中的重大空白。为此,我们提出MicroG-4M,首个面向微重力环境下人类活动时空与语义理解的基准数据集。该数据集基于真实太空任务与影视模拟构建,包含4,759段视频剪辑,涵盖50类动作、1,238条富含上下文的描述文本,以及超过7,000个关于宇航员活动与场景理解的问答对。MicroG-4M支持三项核心任务:细粒度多标签动作识别、时间视频描述生成与视觉问答,可全面评估微重力场景下的空间定位与语义推理能力。我们使用前沿模型建立了基线,并公开所有数据、标注与代码,地址为https://github.com/LEI-QI-233/HAR-in-Space。
原文摘要 · Abstract (English)
Despite substantial progress in video understanding, most existing datasets are limited to Earth's gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-robust video understanding in safety-critical space applications. To address this, we introduce MicroG-4M, the first benchmark for spatio-temporal and semantic understanding of human activities in microgravity. Constructed from real-world space missions and cinematic simulations, the dataset includes 4,759 clips covering 50 actions, 1,238 context-rich captions, and over 7,000 question-answer pairs on astronaut activities and scene understanding. MicroG-4M supports three core tasks: fine-grained multi-label action recognition, temporal video captioning, and visual question answering, enabling a comprehensive evaluation of both spatial localization and semantic reasoning in microgravity contexts. We establish baselines using state-of-the-art models. All data, annotations, and code are available at https://github.com/LEI-QI-233/HAR-in-Space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。