arXiv:2510.26580cs.CV2025-10

让AI在陌生场景中零样本理解环境,靠视觉与语言对齐

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

  • 用预训练视觉模型和大语言模型对齐图像与文字语义
  • 动态推理模块提升复杂场景下18%的识别准确率
  • 适合部署在无标注数据的动态真实环境中

在真实世界中,人工智能系统常面临无标签数据的未知场景,传统场景理解模型难以泛化。本文提出一种动态上下文感知场景推理框架,利用视觉-语言对齐应对零样本现实场景。通过结合预训练视觉变换器与大语言模型,将视觉语义与自然语言描述对齐,增强上下文理解能力。动态推理模块基于语言先验,融合全局场景线索与物体级交互,优化预测结果。在COCO、Visual Genome和Open Images等零样本基准上实验表明,该方法在复杂且未见环境中相比基线模型准确率提升最高达18%。在模糊或杂乱场景中也表现出鲁棒性,得益于视觉与语言的协同融合。该框架为动态真实场景中的上下文感知推理提供了可扩展、可解释的解决方案,推动了零样本泛化能力的发展。

原文摘要 · Abstract (English)

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limits the deployment of vision-based applications in dynamic, unstructured settings. This work introduces a Dynamic Context-Aware Scene Reasoning framework that leverages Vision-Language Alignment to address zero-shot real-world scenarios. The goal is to enable intelligent systems to infer and adapt to new environments without prior task-specific training. The proposed approach integrates pre-trained vision transformers and large language models to align visual semantics with natural language descriptions, enhancing contextual comprehension. A dynamic reasoning module refines predictions by combining global scene cues and object-level interactions guided by linguistic priors. Extensive experiments on zero-shot benchmarks such as COCO, Visual Genome, and Open Images demonstrate up to 18% improvement in scene understanding accuracy over baseline models in complex and unseen environments. Results also show robust performance in ambiguous or cluttered scenes due to the synergistic fusion of vision and language. This framework offers a scalable and interpretable approach for context-aware reasoning, advancing zero-shot generalization in dynamic real-world settings.

零样本学习视觉语言对齐动态推理场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。