arXiv:2506.03155cs.LGcs.AI2025-06被引 4

跨域多模态知识融合,让AI更好解决现实问题

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

  • 构建四层框架,系统化整合不同领域数据
  • 解决跨域数据因结构差异导致的知识对齐难题
  • 适合需要融合多种传感器或系统数据的工程场景

人工智能的发展推动了数字与物理世界间的应用融合。由于物理环境复杂,单一信息采集方式难以覆盖,需融合来自传感器、设备、系统及人员等多源的多模态数据来解决实际问题。然而,为每个问题重新部署资源采集原始数据既不现实也不可持续。因此,当目标领域数据不足时,关键在于融合其他已有领域的多模态知识。我们称之为跨域知识融合。现有研究集中于单域内的多模态数据融合,假设不同数据集的知识内在对齐;但在跨域场景中,这一假设可能不成立。本文首次形式化定义跨域多模态数据融合问题,分析其独特挑战、与单域融合的差异及优势。提出一个包含领域、链接、模型、数据四层的框架,回答‘融什么’、‘为何可融’、‘如何融’三个核心问题。该框架能有效设计跨域多模态数据融合方案,助力解决真实世界问题。

原文摘要 · Abstract (English)

The proliferation of artificial intelligence has enabled a diversity of applications that bridge the gap between digital and physical worlds. As physical environments are too complex to model through a single information acquisition approach, it is crucial to fuse multimodal data generated by different sources, such as sensors, devices, systems, and people, to solve a problem in the real world. Unfortunately, it is neither applicable nor sustainable to deploy new resources to collect original data from scratch for every problem. Thus, when data is inadequate in the domain of problem, it is vital to fuse knowledge from multimodal data that is already available in other domains. We call this cross-domain knowledge fusion. Existing research focus on fusing multimodal data in a single domain, supposing the knowledge from different datasets is intrinsically aligned; however, this assumption may not hold in the scenarios of cross-domain knowledge fusion. In this paper, we formally define the cross-domain multimodal data fusion problem, discussing its unique challenges, differences and advantages beyond data fusion in a single domain. We propose a four-layer framework, consisting of Domains, Links, Models and Data layers, answering three key questions:"what to fuse", "why can be fused", and "how to fuse". The Domains Layer selects relevant data from different domains for a given problem. The Links Layer reveals the philosophy of knowledge alignment beyond specific model structures. The Models Layer provides two knowledge fusion paradigms based on the fundamental mechanisms for processing data. The Data Layer turns data of different structures, resolutions, scales and distributions into a consistent representation that can be fed into an AI model. With this framework, we can design solutions that fuse cross-domain multimodal data effectively for solving real-world problems.

跨域融合多模态知识融合现实问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。