提出可区分共性与个性信息的多模态数据融合方法
Personalized Coupled Tensor Decomposition for Multimodal Data Fusion: Uniqueness and Algorithms
- 将每份数据拆分为共性与个性化两部分,分别建模
- 在真实数据上比现有方法更准确地还原隐藏结构
- 适合处理异构多源数据,如跨模态医疗或社交分析
耦合张量分解(CTD)通过关联不同数据集的因子实现数据融合。然而,现有方法未充分应对数据融合中的关键挑战:一是数据常为同一现象的不同‘视角’(多模态性);二是每份数据可能包含独有的个性化信息,与其他数据无共享因子。本文提出一种个性化耦合张量分解框架,针对每份数据,将其表示为两部分之和:一部分通过多线性测量模型与公共张量关联,另一部分仅属于该数据集。公共与个性成分均假设具有多项式分解形式,此模型推广了多个已有CTD模型。我们给出了分解唯一性的具体与通用条件,易于理解,依赖于各数据集的单模唯一性及测量模型性质。提出了两种算法求解:半代数法与坐标下降优化法。实验表明,该框架在多个真实数据集上优于当前最优方法。
原文摘要 · Abstract (English)
Coupled tensor decompositions (CTDs) perform data fusion by linking factors from different datasets. Although many CTDs have been already proposed, current works do not address important challenges of data fusion, where: 1) the datasets are often heterogeneous, constituting different "views" of a given phenomena (multimodality); and 2) each dataset can contain personalized or dataset-specific information, constituting distinct factors that are not coupled with other datasets. In this work, we introduce a personalized CTD framework tackling these challenges. A flexible model is proposed where each dataset is represented as the sum of two components, one related to a common tensor through a multilinear measurement model, and another specific to each dataset. Both the common and distinct components are assumed to admit a polyadic decomposition. This generalizes several existing CTD models. We provide conditions for specific and generic uniqueness of the decomposition that are easy to interpret. These conditions employ uni-mode uniqueness of different individual datasets and properties of the measurement model. Two algorithms are proposed to compute the common and distinct components: a semi-algebraic one and a coordinate-descent optimization method. Experimental results illustrate the advantage of the proposed framework compared with the state of the art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。