让视觉模型学会像人一样互动理解,提升跨任务泛化能力。
Knowledge Transfer from Interaction Learning
- 通过交互查询和跨模态注意力监督,显式建模视觉理解过程。
- 在TinyImageNet、COCO等任务上提升3.3%和2.4%准确率,零样本迁移增益达9.3%。
- 适合关注知识迁移、跨域泛化与认知对齐的研究者。
当前视觉基础模型(VFMs)在从视觉语言模型(VLMs)迁移知识时面临根本性挑战,尽管VLMs擅长通过统一表征空间建模跨模态交互,但现有VFMs多采用结果导向范式,忽略底层交互过程。这种表征差异阻碍了有效知识迁移并限制了在多样化视觉任务中的泛化能力。我们提出学习交互(LFI)框架,受认知启发,将视觉理解视为一个交互过程。核心洞察在于:捕捉预训练VLM中编码的动态交互模式,可实现更忠实高效的知识迁移。该方法包含两项关键技术:保持层间关系结构的交互查询,以及源自VLM跨模态注意力机制的交互式监督。大量实验表明,该框架在多个基准测试中持续提升性能,在TinyImageNet分类和COCO检测/分割任务上分别获得3.3和1.6mAP/2.4AP的绝对提升,参数开销小且收敛更快。尤其在跨领域设置下表现优异,于PACS和VLCS上实现2.4和9.3的零样本提升。人工评估进一步验证其认知一致性,语义一致度指标优于结果导向方法2.7倍。
原文摘要 · Abstract (English)
Current visual foundation models (VFMs) face a fundamental limitation in transferring knowledge from vision language models (VLMs), while VLMs excel at modeling cross-modal interactions through unified representation spaces, existing VFMs predominantly adopt result-oriented paradigms that neglect the underlying interaction processes. This representational discrepancy hinders effective knowledge transfer and limits generalization across diverse vision tasks. We propose Learning from Interactions (LFI), a cognitive-inspired framework that addresses this gap by explicitly modeling visual understanding as an interactive process. Our key insight is that capturing the dynamic interaction patterns encoded in pre-trained VLMs enables more faithful and efficient knowledge transfer to VFMs. The approach centers on two technical innovations, Interaction Queries, which maintain persistent relational structures across network layers, and interaction-based supervision, derived from the cross-modal attention mechanisms of VLMs. Comprehensive experiments demonstrate consistent improvements across multiple benchmarks, achieving 3.3 and 1.6mAP/2.4AP absolute gains on TinyImageNet classification and COCO detection/segmentation respectively, with minimal parameter overhead and faster convergence. The framework particularly excels in cross-domain settings, delivering 2.4 and 9.3 zero-shot improvements on PACS and VLCS. Human evaluations further confirm its cognitive alignment, outperforming result-oriented methods by 2.7 times in semantic consistency metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。