首个大规模真实360视频数据集,支持多任务学习。
Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
- 用2D分割器+大模型自动标注,分三阶段生成高质量标签。
- 覆盖室内外多种场景,显著提升360视频分割与追踪性能。
- 适合做全景视觉理解、自动驾驶、机器人感知的研究者使用。
360视频以360×180度超大视场角捕捉完整环境,对自动驾驶、机器人等应用中的场景理解(如分割与追踪)至关重要。然而,由于球面特性(如极区严重畸变、内容不连续)导致标注成本高、过程复杂,当前社区缺乏大规模、带标签的真实世界360视频数据集。本文提出Leader360V,首个面向实例分割与追踪的大规模真实360视频数据集。数据涵盖室内、城市、自然及动态户外等多种场景。为实现自动化标注,设计了一套三阶段流水线:初始标注阶段引入语义与畸变感知优化模块,融合多个2D分割器的掩码提议与大语言模型验证的语义标签,生成掩码提示引导SAM2生成畸变感知掩码;自动修正阶段通过重运行SDR或修复水平边界不连续性补全缺失区域;人工修订阶段结合大语言模型与人工标注进一步校验。用户研究与评估证实该流程有效。实验表明,Leader360V可显著提升360视频分割与追踪模型性能,推动更可扩展的全景场景理解发展。
原文摘要 · Abstract (English)
360 video captures the complete surrounding scenes with the ultra-large field of view of 360X180. This makes 360 scene understanding tasks, eg, segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community is, however, impeded by the lack of large-scale, labelled real-world datasets. This is caused by the inherent spherical properties, eg, severe distortion in polar regions, and content discontinuities, rendering the annotation costly yet complex. This paper introduces Leader360V, the first large-scale, labeled real-world 360 video datasets for instance segmentation and tracking. Our datasets enjoy high scene diversity, ranging from indoor and urban settings to natural and dynamic outdoor scenes. To automate annotation, we design an automatic labeling pipeline, which subtly coordinates pre-trained 2D segmentors and large language models to facilitate the labeling. The pipeline operates in three novel stages. Specifically, in the Initial Annotation Phase, we introduce a Semantic- and Distortion-aware Refinement module, which combines object mask proposals from multiple 2D segmentors with LLM-verified semantic labels. These are then converted into mask prompts to guide SAM2 in generating distortion-aware masks for subsequent frames. In the Auto-Refine Annotation Phase, missing or incomplete regions are corrected either by applying the SDR again or resolving the discontinuities near the horizontal borders. The Manual Revision Phase finally incorporates LLMs and human annotators to further refine and validate the annotations. Extensive user studies and evaluations demonstrate the effectiveness of our labeling pipeline. Meanwhile, experiments confirm that Leader360V significantly enhances model performance for 360 video segmentation and tracking, paving the way for more scalable 360 scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。