用轻量稀疏解耦表示提升多模态排除查询的可解释性与效果
Answering Multimodal Exclusion Queries with Lightweight Sparse Disentangled Representations
- 设计固定尺寸的小型稀疏解耦表示,提升可解释性
- 在MSCOCO和概念标题数据集上,AP@10最高提升21%
- 适合需要精准控制和可解释性的多模态检索场景
支持跨模态检索的多模态表示广泛应用,但通常缺乏可解释性,难以说明检索结果依据。现有方法如学习稀疏解耦表示,常依赖文本标记,导致嵌入维度过高。本文提出一种生成更小维度、固定尺寸嵌入的方法,不仅实现解耦,还增强对检索任务的控制力。在MSCOCO和Conceptual Captions基准上的排除查询任务中验证了其有效性。实验表明,该方法优于传统密集模型(如CLIP、BLIP、VISTA),AP@10最高提升11%;也优于稀疏解耦模型(如VDR),AP@10最高提升21%。定性结果进一步证明了解耦表示的可解释性优势。
原文摘要 · Abstract (English)
Multimodal representations that enable cross-modal retrieval are widely used. However, these often lack interpretability making it difficult to explain the retrieved results. Solutions such as learning sparse disentangled representations are typically guided by the text tokens in the data, making the dimensionality of the resulting embeddings very high. We propose an approach that generates smaller dimensionality fixed-size embeddings that are not only disentangled but also offer better control for retrieval tasks. We demonstrate their utility using challenging exclusion queries over MSCOCO and Conceptual Captions benchmarks. Our experiments show that our approach is superior to traditional dense models such as CLIP, BLIP and VISTA (gains up to 11% in AP@10), as well as sparse disentangled models like VDR (gains up to 21% in AP@10). We also present qualitative results to further underline the interpretability of disentangled representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。