构建了覆盖完整场景的稀疏全景RGB-D姿态数据集,减少冗余且可追溯。
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

- 用几何一致性约束筛选视角,生成低冗余全景数据
- 每场景仅25帧即覆盖13类房间,覆盖率达标准方法加误差项
- 支持3D视觉学习中的可审计训练,适合几何一致建模任务
现代3D视觉学习依赖于度量3D资产的观测数据,但现有扫描、网格、点云、模拟和重建结果无法直接提供稀疏、可比且几何一致的全景训练接口。密集轨迹重复邻近视角,源特定渲染策略导致标注异质,稀疏启发式可能遗漏关键区域或引入深度不一致观测。本文提出COVER(基于ERP范围-深度映射的覆盖率导向视角筛选),一种无需训练的ERP视角筛选器,将选定视角的几何信息投影至候选ERP探针,评分增量覆盖并惩罚深度冲突。在受限代理误差下,其贪婪覆盖代理保留标准覆盖率近似行为,最多附加一个可加误差项。基于COVER,构建了CM-EVS(覆盖率优化的度量ERP视角集),包含来自Blender Indoor、HM3D、ScanNet++的1,275个室内场景共36,373帧稀疏全景帧,以及从TartanGround和OB3D重编码的室外全景图,均采用统一格式。每帧含全球面RGB、度量距离深度与校准位姿;室内帧附带逐步溯源日志。平均每场景仅25帧即可覆盖全部13类统一房间类型,保持紧凑的场景级覆盖。实验表明COVER提升了覆盖与冲突间的权衡,使CM-EVS成为稀疏、紧凑、可审计的几何一致全景3D学习资源。
原文摘要 · Abstract (English)
Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a sparse, comparable, and geometry-consistent panoramic training interface. Dense trajectories duplicate nearby views, source-specific rendering policies yield heterogeneous annotations, and sparse heuristics may miss important regions or introduce depth-inconsistent observations. We study how to convert 3D assets into sparse panoramic RGB-D-pose data that preserves complete scene coverage with low redundancy and auditable provenance. We propose COVER (Coverage-Oriented Viewpoint curation with ERP Range-depth warping), a training-free ERP viewpoint curator that projects geometry observed from selected views into candidate ERP probes, scores incremental coverage, and penalizes depth conflicts. Under bounded proxy error, its greedy coverage proxy preserves the standard coverage-style approximation behavior up to an additive error term. Using COVER, we build CM-EVS (Coverage-curated Metric ERP View Set), a panoramic RGB-D-pose dataset with 36,373 curated ERP frames from 1,275 indoor scenes across Blender indoor, HM3D, and ScanNet++, complemented by outdoor panoramas from TartanGround and OB3D re-encoded into the same schema. Each frame provides full-sphere RGB, metric range depth, calibrated pose; COVER-produced indoor frames include per-step provenance logs. With a median of only 25 frames per indoor scene, CM-EVS covers all 13 unified room types while maintaining compact scene-level coverage. Experiments show that COVER improves the coverage-conflict trade-off, making CM-EVS a sparse, compact, and auditable RGB-D-pose resource for geometry-consistent panoramic 3D learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。