让大模型学会跨模态理解,无需训练就能识别新视觉类型。
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis

- 通过合成多种视觉风格图像,让模型学会剥离外观差异抓取语义本质。
- 在6种真实与合成模态上测试,5个模型零样本泛化能力显著提升。
- 适合研究多模态泛化、少样本视觉理解的开发者和研究人员。
尽管大型多模态模型(LMMs)在RGB视觉任务上取得进展,其对未见视觉模态的泛化能力仍基本未被探索。我们认为不同视觉模态只是同一物理世界的不同采样方式,因此有效的泛化需模型具备模态无关的场景语义感知能力及对模态特性的适应能力。为此,我们提出一种训练框架VVM-Tuning,通过模态合成与模态上下文实现该目标。具体而言,从RGB场景中合成多样化外观图像,训练模型将不变语义与可变外观解耦,并将这些外观与语言对齐,实现脱离模态的视觉概念表征。随后在提示中引入模态上下文,通过指令微调帮助模型将外观变化映射回模态相关属性,从而在推理时实现对未见模态的零样本适应。为推动此方向研究,我们构建了VVM-Bench,一个包含6种真实与合成模态的综合性基准,用于评估语义感知与模态理解能力。实验表明,仅在合成模态上训练后,5个测试模型在真实世界及新型合成模态上均表现出一致提升,且无需模态内训练。
原文摘要 · Abstract (English)
Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world. Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at https://github.com/Hunter-Will/VVM-Tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。