仅用文本描述实现跨模态扩展,无需成对数据。
TextME: Bridging Unseen Modalities Through Text Descriptions
- 利用文本描述将多种模态映射到统一语言嵌入空间。
- 在图像、音频、3D、X光、分子等域中保持预训练模型性能。
- 支持未显式对齐模态间的零样本跨模态检索,如音频到图像。
将多模态表示拓展至新模态受限于大规模成对数据集(如文本-图像、文本-音频、文本-3D、文本-分子)的依赖,这类数据在医学影像和分子分析等需专家标注的领域成本高昂且难以获取。我们提出TextME,据我们所知首个纯文本模态扩展框架,将多种模态投影至大型语言模型(LLM)嵌入空间作为统一锚点。该方法利用预训练对比编码器的几何结构,仅通过文本描述即可实现零样本跨模态迁移,无需成对监督。实证验证了图像、视频、音频、3D、X光、分子等领域存在一致的模态差异,表明纯文本训练可有效保留预训练编码器的性能。进一步证明该框架能实现训练中未明确对齐的模态对之间的涌现式跨模态检索(如音频→图像、3D→图像)。这些结果确立了纯文本训练作为模态扩展中成对监督的可行替代方案。
原文摘要 · Abstract (English)
Expanding multimodal representations to novel modalities is constrained by reliance on large-scale paired datasets (e.g., text-image, text-audio, text-3D, text-molecule), which are costly and often infeasible in domains requiring expert annotation such as medical imaging and molecular analysis. We introduce TextME, the first text-only modality expansion framework, to the best of our knowledge, projecting diverse modalities into LLM embedding space as a unified anchor. Our approach exploits the geometric structure of pretrained contrastive encoders to enable zero-shot cross-modal transfer using only text descriptions, without paired supervision. We empirically validate that such consistent modality gaps exist across image, video, audio, 3D, X-ray, and molecular domains, demonstrating that text-only training can preserve substantial performance of pretrained encoders. We further show that our framework enables emergent cross-modal retrieval between modality pairs not explicitly aligned during training (e.g., audio-to-image, 3D-to-image). These results establish text-only training as a practical alternative to paired supervision for modality expansion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。