arXiv:2412.10348cs.CVcs.AI2024-12被引 1

通过细粒度对齐提升视觉区域理解,增强跨模态生成效果

A dual contrastive framework

  • 设计双对比框架,优化潜在空间表示以实现区域级对齐
  • 在多个任务上显著提升区域描述生成性能,改善跨模态对齐质量
  • 适合关注视觉-语言对齐与区域级理解的研究者

当前多模态任务中,模型通常冻结编码器和解码器,仅调整中间层以适应特定目标(如区域描述)。大尺度视觉-语言模型在区域级视觉理解方面面临挑战。尽管空间感知能力有限是已知问题,但粗粒度预训练进一步加剧了潜在表示优化的难度,影响编码器-解码器对齐效果。本文提出AlignCap框架,通过细粒度潜在空间对齐增强区域级理解。方法引入新型潜在特征精炼模块,提升条件化潜在表示,改善区域描述性能;并提出语义空间对齐模块,提升多模态表示质量。同时,在两个模块中创新性地融合对比学习,进一步强化区域描述表现。为缓解空间限制,采用通用目标检测(GOD)作为数据预处理流程,增强区域级空间推理能力。大量实验表明,该方法在多种任务中显著提升区域级描述性能。

原文摘要 · Abstract (English)

In current multimodal tasks, models typically freeze the encoder and decoder while adapting intermediate layers to task-specific goals, such as region captioning. Region-level visual understanding presents significant challenges for large-scale vision-language models. While limited spatial awareness is a known issue, coarse-grained pretraining, in particular, exacerbates the difficulty of optimizing latent representations for effective encoder-decoder alignment. We propose AlignCap, a framework designed to enhance region-level understanding through fine-grained alignment of latent spaces. Our approach introduces a novel latent feature refinement module that enhances conditioned latent space representations to improve region-level captioning performance. We also propose an innovative alignment strategy, the semantic space alignment module, which boosts the quality of multimodal representations. Additionally, we incorporate contrastive learning in a novel manner within both modules to further enhance region-level captioning performance. To address spatial limitations, we employ a General Object Detection (GOD) method as a data preprocessing pipeline that enhances spatial reasoning at the regional level. Extensive experiments demonstrate that our approach significantly improves region-level captioning performance across various tasks

多模态区域描述对比学习视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。