arXiv:2602.00340cs.CVcs.AI2026-02

解决视觉语言模型在未知概念上的跨模态对齐失效问题

Bridging the Semantic Chasm: Synergistic Conceptual Anchoring for Generalized Few-Shot and Zero-Shot OOD Perception

  • 用四个协同模块动态修复跨模态差异
  • 在多个领域少样本和零样本任务中提升1.2%至5.4%精度
  • 适合做开放世界感知与小样本迁移的研究者

本文提出一种名为协同神经代理网络(SynerNet)的创新框架,旨在缓解视觉语言模型在面对分布外(OOD)概念时出现的跨模态对齐退化问题。该框架由视觉感知、语言上下文、命名嵌入和全局协调四个专用计算单元构成,通过结构化的消息传播机制协同校正模态差异。主要贡献包括多智能体潜在空间命名获取框架、增强少样本适应的语义上下文交换算法,以及自适应动态平衡机制。在VISTA-Beyond基准上的实证评估表明,SynerNet在少样本和零样本场景下均取得显著性能提升,精度改善范围为1.2%至5.4%,覆盖多种异构领域。

原文摘要 · Abstract (English)

This manuscript presents a pioneering Synergistic Neural Agents Network (SynerNet) framework designed to mitigate the phenomenon of cross-modal alignment degeneration in Vision-Language Models (VLMs) when encountering Out-of-Distribution (OOD) concepts. Specifically, four specialized computational units - visual perception, linguistic context, nominal embedding, and global coordination - collaboratively rectify modality disparities via a structured message-propagation protocol. The principal contributions encompass a multi-agent latent space nomenclature acquisition framework, a semantic context-interchange algorithm for enhanced few-shot adaptation, and an adaptive dynamic equilibrium mechanism. Empirical evaluations conducted on the VISTA-Beyond benchmark demonstrate that SynerNet yields substantial performance augmentations in both few-shot and zero-shot scenarios, exhibiting precision improvements ranging from 1.2% to 5.4% across a diverse array of domains.

视觉语言模型少样本学习分布外泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。