arXiv:2608.01852cs.DBcs.IR2026-08

为3D数字孪生平台添加多模态嵌入,实现语义与相似性混合检索

Multimodal Embeddings for 3D Similarity Search in Semantic Web-of-Things Digital-Twin Platforms

  • 将3D点云、时间属性和语义标签统一编码为向量嵌入
  • 在S3DIS数据集上验证了图过滤可有效缩小搜索范围
  • 适合需要融合结构查询与视觉相似性的工业物联网场景

语义物联网(SWoT)平台将物理基础设施建模为基于领域本体的知识图谱,支持丰富的结构化与逻辑查询。然而,其缺乏表达相似性的原生机制,这在电信基础设施和工业物联网等3D数字孪生领域构成关键瓶颈——查询需结合本体约束与异构、时序演化的场景数据的多模态相似性搜索。本文提出一种框架,在SWoT平台中引入多模态嵌入层:将包含3D点云、时间属性和语义标签的本体类型实体编码为潜在向量表示,并存储于知识图谱旁,支持混合的图-向量查询,即结合图谱过滤与相似性检索。该框架在Orange Research的Thing'in平台及Clock-G时序图数据库上实现,对S3DIS数据集的可行性评估表明,图过滤在时间与关系约束下能有效缩小搜索池;通用预训练编码器生成的表征足以支持相似性检索,并可作为下游预测任务的初步编码步骤。

原文摘要 · Abstract (English)

Semantic Web of Things (SWoT) platforms model physical infrastructure as knowledge graphs typed against domain ontologies, enabling expressive structural and logical queries. However, they lack native mechanisms to express similarity beyond strict ontological equivalence, which represents a critical gap for 3D digital twins in domains such as telecom infrastructure and industrial IoT, where queries must combine ontological constraints with multimodal similarity search over heterogeneous, temporally-evolving scene data. We propose a framework that extends SWoT platforms with a multimodal embedding layer: ontology-typed entities comprising 3D point clouds, temporal attributes, and semantic labels are encoded into latent vector representations stored alongside the knowledge graph, enabling hybrid ontology-vector queries that combine graph-based filtering with similarity search. Implemented on Orange Research's Thing'in platform with the Clock-G temporal graph database, a feasibility evaluation on S3DIS demonstrates that graph filtering effectively restricts the search pool under temporal and relational constraints, and that general-purpose pretrained encoders produce representations sufficient for similarity retrieval and as a preliminary encoding step for downstream predictive tasks.

数字孪生多模态嵌入语义网3D检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。