将2D图像理解迁移到3D空间,为室内安全设备识别提供高效智能方案。
INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety
- 通过注册的RGB-D数据,将2D视觉模型输出映射到3D点云空间。
- 在斯坦福2D-3D-S数据集上实现7个共享类的点级标注准确率,15类关键安全设备检测灵敏度高。
- 生成符合标准的场景图,压缩后可在15秒内传输,适合应急通信网络部署。
室内环境缺乏如室外GPS般的位置智能基础设施;突发事件中,救援人员常无法获取可机器读取的安全设施地图。现有3D语义分割研究面临两大障碍:标注数据稀缺,以及原始点云方法对小型关键设施识别能力差。本文提出INSIGHT,一种无需目标域标注的零样本管道,通过注册的RGB-D数据将2D图像理解投影至3D度量空间。两个可互换的视觉模块共享同一3D后端:基于SAM3的基础模型模块支持文本提示分割,传统计算机视觉模块(开放集检测、VQA、OCR)的中间输出可独立检视。在斯坦福2D-3D-S的全部七个子区域(共70,496张图像)上评估,该管道生成与Pointcept兼容的标注点云及符合ISO 19164标准的场景图,压缩比达约10⁴倍;角色过滤后的数据包在1 Mbps、FirstNet Band 14条件下传输时间少于15秒。报告了7个共用类别的点级标注准确率,15类未出现在公开3D基准中的关键安全设备检测灵敏度,以及代码受限的可部署估算结果,表明2D到3D语义迁移有效缓解标注数据瓶颈,场景图具备足够紧凑性以支持现场部署。
原文摘要 · Abstract (English)
Indoor environments lack the spatial intelligence infrastructure that GPS provides outdoors; first responders arriving at unfamiliar buildings typically have no machine-readable map of safety equipment. Prior work on 3D semantic segmentation for public safety identified two barriers: scarcity of labeled indoor training data and poor recognition of small safety-critical features by native point-cloud methods. This paper presents INSIGHT, a zero-target-domain-annotation pipeline that projects 2D image understanding into 3D metric space via registered RGB-D data. Two interchangeable vision stacks share a common 3D back end: a SAM3 foundation-model stack for text-prompted segmentation, and a traditional CV stack (open-set detection, VQA, OCR) whose intermediate outputs are independently inspectable. Evaluated on all seven subareas of Stanford 2D-3D-S (70{,}496 images), the pipeline produces Pointcept-schema-compatible labeled point clouds and ISO~19164-compliant scene graphs with ${\sim}10^{4}{\times}$ compression; role-filtered payloads transmit in ${<}15$\,s at 1\,Mbps over FirstNet Band~14. We report per-point labeling accuracy on 7 shared classes, detection sensitivity for 15 safety-critical classes absent from public 3D benchmarks alongside code-capped deployable estimates, and inter-pipeline complementarity, demonstrating that 2D-to-3D semantic transfer addresses the labeled-data bottleneck while scene graphs provide building intelligence compact enough for field deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。