arXiv:2503.17116cs.MMcs.AI2025-03被引 18

首个无遮挡的双视角多模态视频数据集,支持第一人称与第三人称同步研究。

The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding

  • 采集10人第一视角+5个固定摄像头第三人视角,时间对齐的多模态数据
  • 超600小时4K高清视频,50帧/秒,无任何隐私遮挡处理
  • 适合动作识别、人机交互、多视角理解等方向研究者使用

近年来,第一人称视频受到广泛关注,但多数现有数据集仅包含单一视角。本文提出CASTLE 2024数据集,是一个多模态数据集,包含来自15个时间对齐源的自我中心(第一人称)和外部中心(第三人称)视频与音频,以及其它传感器流和辅助数据。数据由志愿者在固定地点连续四天录制,涵盖10名参与者的视角,另有5个固定摄像机提供第三人称视角。整个数据集包含超过600小时的4K超高清视频,采样率为50帧每秒。与现有数据集不同,CASTLE 2024不包含任何形式的部分遮蔽,如模糊人脸或失真音频。数据集可通过https://castle-dataset.github.io/获取。

原文摘要 · Abstract (English)

Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a fixed location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.

多模态第一人称视频数据集行为理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。