arXiv:2606.18952cs.CV2026-06

首个真实采集的单光子感知多任务基准,助力复杂环境3D重建与语义理解。

SP-TransientBench: A Real-Captured Single Photon Perception Benchmark

论文配图:SP-TransientBench: A Real-Captured Single Photon Perception Benchmark
图 1 · 摘自论文原文
  • 构建真实世界单光子激光雷达数据集,含10个场景、10297视角及时间-飞行直方图。
  • 支持深度估计、多视角重建与3D语义理解,提供13类语义标注和标准化评估协议。
  • 填补真实数据空白,适合研究单光子感知、低光照3D视觉的学者使用。

基于单光子雪崩二极管(SPAD)的单光子激光雷达(SPL)具备极高灵敏度,可在光子稀缺场景下实现时间分辨的光子测量,为主动三维感知提供独特潜力。然而,真实世界的单光子感知仍面临独特测量噪声与复杂的多回波瞬态现象挑战,共同加剧几何重建与语义场景理解的难度。尽管对SPAD传感的兴趣日益增长,现有研究大多局限于模拟数据或小规模受控采集。因此,针对真实世界单光子感知在深度估计、多视角重建与3D语义理解等方面的系统性评估仍严重不足。为弥合这一差距,我们提出SP-TransientBench(STB),一个真实采集的多任务单光子感知基准。STB包含10个多样化场景与10,297个视角,采用固态单光子激光雷达以$256\times192$分辨率采集,每个视角提供完整的时间飞行直方图、多回波行为、标准化元数据与校准相机位姿,支持多视角评估。我们进一步为部分场景提供13类3D语义标注。通过为每项任务设计专用数据划分与评估协议,STB实现了真实世界单光子感知在多个3D视觉问题上的可重复、一致性基准测试。数据集与代码将在论文接受后发布。

原文摘要 · Abstract (English)

Single-photon LiDAR (SPL) based on single-photon avalanche diode (SPAD) sensing enables time-resolved photon measurements with extreme sensitivity, offering unique potential for active 3D perception in photon-starved scenarios.However, real-world single photon perception remains fundamentally challenging due to unique measurement noise and complex multi-return transient phenomena, which jointly complicate geometric reconstruction and semantic scene understanding. Despite growing interest in SPAD-based sensing, existing studies are largely limited to simulated data or small-scale controlled captures. As a result, systematic evaluation of real-world single photon perception across depth estimation, multi-view reconstruction, and 3D semantic understanding remains underexplored. To bridge this gap, we introduce SP-TransientBench (STB), a real-captured multi-task benchmark for single photon perception. SP-TransientBenc comprises 10 diverse scenes and 10,297 views captured using a solid-state single-photon LiDAR at $256\times192$ resolution. Each view provides full time-of-flight histograms with multi-return behavior,standardized metadata, and calibrated camera poses for multi-view evaluation. We further provide 13-class 3D semantic annotations for selected scenes. By providing dedicated data splits and evaluation protocols for each task, STB enables consistent and reproducible benchmarking of real-world single photon perception across multiple 3D vision problems. The dataset and code will be released upon acceptance.

单光子感知3D重建语义理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。