用大模型提升图文对齐,解决文本图像信息熵差异问题。
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- 用大模型生成多义描述增强文本信息熵,改善跨模态对齐。
- 通过超图适配器修复语义同义匹配错误,提升检索准确率。
- 在Flickr30K和MS-COCO上实现图文检索显著提升,适合多模态研究者。
图文对齐是多媒体内容理解的基础挑战,有效建模跨模态语义对应关系可通过联合嵌入空间优化显著提升检索系统性能。由于文本与图像间存在固有的信息熵差异,传统方法常导致两模态互检失衡。为此,我们提出利用大语言模型(LLM)的开放语义知识填补熵差,并重现人类在该任务中的对齐能力。我们的熵增强对齐采用两步策略:1)设计一种不依赖特定任务领域显式知识的新提示模板,通过LLM增强文本模态的多义性描述,从而提高文本相对于视觉模态的信息熵;2)使用超图适配器在文本与图像模态间构建多边连接,可修正同一固定嵌入空间中同义语义的正负匹配错误,同时通过降维再映射回原维度,降低开放语义熵带来的噪声。在Flickr30K和MS-COCO基准上的综合评估验证了所提开放语义超图适配器(OS-HGAdapter)的优越性,相较现有方法在文本到图像检索上提升16.8%,图像到文本检索上提升40.1%,并在语义对齐任务中创下新纪录。
原文摘要 · Abstract (English)
Text-image alignment constitutes a foundational challenge in multimedia content understanding, where effective modeling of cross-modal semantic correspondences critically enhances retrieval system performance through joint embedding space optimization. Given the inherent difference in information entropy between texts and images, conventional approaches often show an imbalance in the mutual retrieval of these two modalities. To address this particular challenge, we propose to use the open semantic knowledge of Large Language Model (LLM) to fill for the entropy gap and reproduce the alignment ability of humans in these tasks. Our entropy-enhancing alignment is achieved through a two-step process: 1) a new prompt template that does not rely on explicit knowledge in the task domain is designed to use LLM to enhance the polysemy description of the text modality. By analogy, the information entropy of the text modality relative to the visual modality is increased; 2) A hypergraph adapter is used to construct multilateral connections between the text and image modalities, which can correct the positive and negative matching errors for synonymous semantics in the same fixed embedding space, whilst reducing the noise caused by open semantic entropy by mapping the reduced dimensions back to the original dimensions. Comprehensive evaluations on the Flickr30K and MS-COCO benchmarks validate the superiority of our Open Semantic Hypergraph Adapter (OS-HGAdapter), showcasing 16.8\% (text-to-image) and 40.1\% (image-to-text) cross-modal retrieval gains over existing methods while establishing new state-of-the-art performance in semantic alignment tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。