arXiv:2409.18868cs.CLcs.AI2024-09

CLIP比纯文本模型更准确捕捉物体数量差异。

Individuation in Neural Models with and without Visual Grounding

  • 对比视觉-语言与纯文本模型,分析个体化信息编码能力。
  • CLIP能更好区分物体数量与聚合状态的细微差异。
  • 结果符合语言学与认知科学中的个体化层级理论,适合认知研究者参考。

我们比较了视觉-语言模型CLIP与两种纯文本模型FastText、SBERT在个体化信息编码上的差异。研究其对物质实体、颗粒集合及不同数量物体的潜在表示。结果显示,CLIP嵌入能更准确地捕捉个体化程度的量化差异,优于仅基于文本训练的模型。进一步地,从CLIP嵌入中推导出的个体化层级,与语言学和认知科学中提出的层级结构高度一致。

原文摘要 · Abstract (English)

We show differences between a language-and-vision model CLIP and two text-only models - FastText and SBERT - when it comes to the encoding of individuation information. We study latent representations that CLIP provides for substrates, granular aggregates, and various numbers of objects. We demonstrate that CLIP embeddings capture quantitative differences in individuation better than models trained on text-only data. Moreover, the individuation hierarchy we deduce from the CLIP embeddings agrees with the hierarchies proposed in linguistics and cognitive science.

多模态个体化语义表征认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。