统一多模态实体定义,让知识图谱更懂图文音视频融合语义。
A Pattern to Align Them All: Integrating Different Modalities to Define Multi-Modal Entities
- 提出分离实体语义与媒体表现的抽象建模方法
- 支持跨领域多模态数据统一整合,提升知识图谱兼容性
- 适合医疗、数字人文等需融合多源信息的场景
人类智能依赖于对多种感官输入的推理与整合,这推动了多模态信息在知识图谱中的建模研究。多模态知识图谱通过为实体关联文本、图像、音频、视频等多种媒介表示,以传达其语义。尽管该领域受关注度日益上升,但对模态的定义尚无共识,常由应用领域决定。本文提出一种新型本体设计模式,将实体及其语义表达(可在不同媒介中呈现)与其物理信息实体实现相分离。该抽象模型旨在促进现有异构多模态本体的协调与集成,对医疗、数字人文等多个领域的智能应用至关重要。
原文摘要 · Abstract (English)
The ability to reason with and integrate different sensory inputs is the foundation underpinning human intelligence and it is the reason for the growing interest in modelling multi-modal information within Knowledge Graphs. Multi-Modal Knowledge Graphs extend traditional Knowledge Graphs by associating an entity with its possible modal representations, including text, images, audio, and videos, all of which are used to convey the semantics of the entity. Despite the increasing attention that Multi-Modal Knowledge Graphs have received, there is a lack of consensus about the definitions and modelling of modalities, whose definition is often determined by application domains. In this paper, we propose a novel ontology design pattern that captures the separation of concerns between an entity (and the information it conveys), whose semantics can have different manifestations across different media, and its realisation in terms of a physical information entity. By introducing this abstract model, we aim to facilitate the harmonisation and integration of different existing multi-modal ontologies which is crucial for many intelligent applications across different domains spanning from medicine to digital humanities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。