构建多模态数据集EnvoDat,提升机器人在复杂环境下的感知能力。
EnvoDat: A Large-Scale Multisensory Dataset for Robotic Spatial Awareness and Semantic Reasoning in Heterogeneous Environments
- 在13个场景采集26段序列,融合10种传感器模态。
- 含超89万标注,覆盖82类物体与地形,支持高精度语义推理。
- 适配SLAM与多模态模型训练,尤其适合极端环境研究者。
为提升机器人在多样化真实场景下的自主性效率,高质量异构数据集至关重要。现有基准主要聚焦城市道路场景,对地下隧道、自然田野及现代室内等多退化、高植被、动态且特征稀疏的环境覆盖不足。为此,我们推出EnvoDat——一个大规模、多模态数据集,涵盖高光照、雾、雨、全天候零可见度等复杂条件。整体包含13个场景的26段序列,10种传感模态,数据总量超1.9TB,提供超过89,000个细粒度多边形标注,覆盖82类物体与地形。我们对数据进行多种格式后处理,支持SLAM算法、监督学习及多模态视觉模型微调。该数据集助力在极端挑战性环境中实现环境鲁棒的机器人自主。数据与资源可通过https://linusnep.github.io/EnvoDat/获取。
原文摘要 · Abstract (English)
To ensure the efficiency of robot autonomy under diverse real-world conditions, a high-quality heterogeneous dataset is essential to benchmark the operating algorithms' performance and robustness. Current benchmarks predominantly focus on urban terrains, specifically for on-road autonomous driving, leaving multi-degraded, densely vegetated, dynamic and feature-sparse environments, such as underground tunnels, natural fields, and modern indoor spaces underrepresented. To fill this gap, we introduce EnvoDat, a large-scale, multi-modal dataset collected in diverse environments and conditions, including high illumination, fog, rain, and zero visibility at different times of the day. Overall, EnvoDat contains 26 sequences from 13 scenes, 10 sensing modalities, over 1.9TB of data, and over 89K fine-grained polygon-based annotations for more than 82 object and terrain classes. We post-processed EnvoDat in different formats that support benchmarking SLAM and supervised learning algorithms, and fine-tuning multimodal vision models. With EnvoDat, we contribute to environment-resilient robotic autonomy in areas where the conditions are extremely challenging. The datasets and other relevant resources can be accessed through https://linusnep.github.io/EnvoDat/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。