arXiv:2601.19640cs.CV2026-01

面向城市治理的低空视觉系统,聚焦管理需求构建多模态基准与推理框架。

Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework

  • 构建面向管理需求的低空感知多模态数据集GovLA-10K
  • 提出无需微调的隐式协同视觉语言推理框架GovLA-Reasoner
  • 通过空间感知适配器保留细粒度空间信息,提升治理理解能力

低空视觉系统正成为智慧城市建设的关键基础设施。然而,现有以物体为中心的感知范式和松耦合的视觉-语言流程难以满足真实城市管理中对异常行为理解的需求。为此,我们提出首个面向管理任务的低空智能多模态基准GovLA-10K,以及专为治理感知设计的统一视觉-语言推理框架GovLA-Reasoner。不同于以往全面标注可见物体的研究,GovLA-10K围绕与实际管理直接相关的功能显著目标进行设计,并提供基于观测结果的可操作管理建议。为有效协调细粒度视觉定位与高层上下文语言推理,GovLA-Reasoner引入高效的时空感知适配器(SGA),隐式协调视觉检测器与大语言模型(LLM)之间的判别性表示共享。与侧重全局嵌入对齐的现有适配器不同,本SGA专门用于压缩和聚合多流的定位感知表示,从而在保留细粒度空间线索的同时,实现其向语言推理过程的有效融合。大量实验表明,该方法在不微调任何任务特定组件的前提下显著提升性能。我们认为本工作为未来面向管理的低空视觉-语言系统研究提供了新视角与基础。代码与数据集将在进一步整理后公开。

原文摘要 · Abstract (English)

Low-altitude vision systems are becoming a critical infrastructure for smart city governance. However, existing object-centric perception paradigms and loosely coupled vision-language pipelines are still difficult to support management-oriented anomaly understanding required in real-world urban governance. To bridge this gap, we introduce GovLA-10K, the first management-oriented multi-modal benchmark for low-altitude intelligence, along with GovLA-Reasoner, a unified vision-language reasoning framework tailored for governance-aware aerial perception. Unlike existing studies that aim to exhaustively annotate all visible objects, GovLA-10K is deliberately designed around functionally salient targets that directly correspond to practical management needs, and further provides actionable management suggestions grounded in these observations. To effectively coordinate the fine-grained visual grounding with high-level contextual language reasoning, GovLA-Reasoner introduces an efficient Spatially-aware Grounding Adapter (SGA) that implicitly coordinates discriminative representation sharing between the visual detector and the large language model (LLM). Different from existing adapters that primarily focus on global embedding alignment, our SGA is specifically designed to compress and aggregate multi-stream grounding-aware representations, thereby preserving fine-grained spatial cues while enabling their effective integration into the language reasoning process. Extensive experiments indicate that our GovLA-Reasoner effectively improves performance while avoiding the need of fine-tuning for any task-specific individual components. We believe our work offers a new perspective and foundation for future studies on management-aware low-altitude vision-language systems. The code and dataset will be publicly released after further organization.

低空感知视觉语言城市管理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。