让大模型无视图像风格差异,精准理解语义。
Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models
- 分离图像内容与风格,避免风格干扰理解
- 跨风格任务表现优于现有方法,尤其在少样本场景
- 适合需要稳定视觉理解的多模态应用
大型视觉语言模型(LVLMs)面临‘风格陷阱’挑战,难以在不同视觉风格下保持稳健的语义理解,尤其在上下文学习(ICL)中表现不佳。现有方法常无法有效解耦风格与内容,限制泛化能力。为此,我们提出语义保持跨风格视觉推理框架SP-CSVR,包含跨风格特征编码器(CSFE)实现风格-内容解耦、语义对齐上下文解码器(SAICD)实现高效少样本风格适应,以及基于多任务对比学习的自适应语义一致性模块(ASCM),强制跨风格语义不变性。在多风格数据集上的大量实验表明,SP-CSVR在视觉描述生成、视觉问答和上下文风格适应任务上均达到当前最优性能。消融实验与泛化分析证实其在鲁棒性、泛化性和效率方面的有效性。
原文摘要 · Abstract (English)
The "style trap" poses a significant challenge for Large Vision-Language Models (LVLMs), hindering robust semantic understanding across diverse visual styles, especially in in-context learning (ICL). Existing methods often fail to effectively decouple style from content, hindering generalization. To address this, we propose the Semantic-Preserving Cross-Style Visual Reasoner (SP-CSVR), a novel framework for stable semantic understanding and adaptive cross-style visual reasoning. SP-CSVR integrates a Cross-Style Feature Encoder (CSFE) for style-content disentanglement, a Semantic-Aligned In-Context Decoder (SAICD) for efficient few-shot style adaptation, and an Adaptive Semantic Consistency Module (ASCM) employing multi-task contrastive learning to enforce cross-style semantic invariance. Extensive experiments on a challenging multi-style dataset demonstrate SP-CSVR's state-of-the-art performance across visual captioning, visual question answering, and in-context style adaptation. Comprehensive evaluations, including ablation studies and generalization analysis, confirm SP-CSVR's efficacy in enhancing robustness, generalization, and efficiency across diverse visual styles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。