用多模态大模型让点云理解建筑构件功能,支持自然语言交互。
From Geometric Labels to Semantic Understanding of Indoor Building Components Using Multimodal Large Language Models
- 基于点云与指令生成响应,融合几何与语义信息。
- 在三类任务上分别达到88.00%、65.10%、68.14%准确率。
- 适合智能运维、建筑信息建模等需语义理解的场景。
基于点云的室内建筑构件理解已成为设施运维的重要技术支撑。然而现有方法仅输出离散标签,缺乏对构件功能的解释与自然语言交互能力。本文提出Building-MLLM,一种以点云为中心的多模态大语言模型,通过建模点云与指令,在简单识别、复杂描述和多工程问答任务中生成响应。该模型通过四种领域专用机制解决语义集中问题:点信息增强器提升任务相关语义,几何保持正则化防止几何失真,固定文本前缀稳定领域表达,多维LoRA平衡识别与推理。开发了多约束渐进式指令生成引擎,构建合成点云-文本数据集,包含4198个对象、37,782个指令跟随对及47个类别。实验表明,Building-MLLM在三类任务上分别达到88.00%、65.10%、68.14%的性能,展现了优越的室内构件语言理解能力,并在其他真实数据集上表现出初步泛化性。
原文摘要 · Abstract (English)
Point cloud-based understanding has become an important enabler for facility operation and maintenance involving indoor building components. However, existing methods output only discrete labels without explaining component functions or natural language interactions. This paper proposes Building-MLLM, a point cloud-centered multimodal large language model (MLLM) for indoor components, which models point clouds and instructions to generate responses across Simple Recognition, Complex Captioning, and Multi-Engineering Question Answering tasks. Building-MLLM addresses semantic concentration through four domain-specific mechanisms: Point Information Enhancer for task-relevant semantics, Geometry-Preserving Regularization preventing geometric erosion, fixed textual prefix for domain stabilization, and multi-dimensional LoRA balancing recognition with reasoning. A multi-constraint progressive instruction-generation engine is developed to compile a synthetic point cloud-text dataset with 4198 objects, 37,782 instruction-following pairs, and 47 categories. Experiments show that Building-MLLM achieves 88.00%, 65.10%, and 68.14% on the three task types, respectively, demonstrating superior indoor component language understanding and providing initial generalizability in transfer inference on other real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。