arXiv:2412.13847cs.AIcs.LG2024-12

用共享概念空间实现跨模态高效学习,像人一样举一反三。

A Concept-Centric Approach to Multi-Modality Learning

  • 构建与模态无关的概念空间,统一抽象知识表示
  • 新模态接入时收敛更快,训练量减少且无需微调
  • 适合需要快速适配新数据模态的系统设计

人类能通过一致的世界认知,在不同模态间高效获取并迁移知识。受此启发,我们提出一种以概念为中心的多模态学习框架:基于一个与模态无关的概念空间,捕获结构化、抽象的知识;配合一组模态特定的投影模型,将原始输入映射到该共享空间。概念空间独立于具体模态,作为通用知识库。一旦习得,即可加速新模态的适应,因投影模型只需对齐已有概念表示,而非从头学习。实验验证该框架收敛更快。其模块化设计支持新模态无缝集成——投影模型可独立训练,但输出统一于共享概念空间。我们在两个典型下游任务上评估,虽未进行任务优化,仍取得相当性能,仅需更小训练开销,无需任务微调,推理全程在可解释的概念空间中完成。结果表明,该方法为构建更贴近人类认知的学习系统提供了新方向。

原文摘要 · Abstract (English)

Humans possess a remarkable ability to acquire knowledge efficiently and apply it across diverse modalities through a coherent and shared understanding of the world. Inspired by this cognitive capability, we introduce a concept-centric multi-modality learning framework built around a modality-agnostic concept space that captures structured, abstract knowledge, alongside a set of modality-specific projection models that map raw inputs onto this shared space. The concept space is decoupled from any specific modality and serves as a repository of universally applicable knowledge. Once learned, the knowledge embedded in the concept space enables more efficient adaptation to new modalities, as projection models can align with existing conceptual representations rather than learning from scratch. This efficiency is empirically validated in our experiments, where the proposed framework exhibits faster convergence compared to baseline models. In addition, the framework's modular design supports seamless integration of new modalities, since projection models are trained independently yet produce unified outputs within the shared concept space. We evaluate the framework on two representative downstream tasks. While the focus is not on task-specific optimization, the framework attains comparable results with a smaller training footprint, no task-specific fine-tuning, and inference performed entirely within a shared space of learned concepts that offers interpretability. These findings point toward a promising direction for developing learning systems that operate in a manner more consistent with human cognitive processes.

多模态学习概念空间跨模态迁移可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。