提出新方法检测跨模态否定,提升视觉语言模型的语义理解能力。
Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

- 设计交叉模态注意力架构,显式建模视觉与语言间的依赖关系。
- 在3222对政治视频文本上测试,性能比单模态基线最高提升7.03% F1。
- 发现视觉否定依赖语言上下文,适合多模态语义理解研究者。
当前多模态系统在检测跨模态的高阶语义概念(如否定)方面仍面临挑战。本文将其视为基础表征学习问题,首次证明否定在标准视觉-语言模型(VLMs)的隐空间中既非线性也非非线性可分。实验表明,预训练嵌入主要编码模态特异性特征,缺乏通用的否定信号。为此,我们提出一种新型跨模态注意力架构,显式建模模态间依赖关系,在3,222对通过Qwen2.5-VL自动标注的政治视频-文本对上,相较单模态基线实现最高达+7.03%的F1提升。分析揭示关键不对称性:文本否定常独立出现,而视觉否定在语义上依赖语言上下文。结合自监督视频表征JEPA2,进一步推进了时间维度否定建模。本工作为构建鲁棒、语义对齐的多模态表征提供了新方法与洞见。
原文摘要 · Abstract (English)
Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this as a fundamental representation learning problem, providing the first evidence that negation does not form a linearly or non-linearly separable class in the latent spaces of standard vision-language models (VLMs). We demonstrate that pretrained embeddings primarily encode modality-specific features, lacking a generalizable negation signal. To overcome this, we propose a novel cross-modal attention architecture that explicitly models inter-modal dependencies, achieving performance gains of up to +7.03% F1 over unimodal baselines. Our analysis reveals a key asymmetry: while textual negation often appears independently, visual negation is semantically dependent on linguistic context, a finding validated through our statistical analysis of 3,222 political video-text pairs automatically annotated via \textsc{Qwen2.5-VL}. By combining this analysis with self-supervised video representations (JEPA2), we advance the modeling of temporal negation. This work provides new methods and insights for learning robust, semantically-aligned representations in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。