用视觉定位提升零样本公交视频分析,降低云端计算成本。
GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning

- 边缘端追踪车门状态并提取乘客片段,云端仅处理定位后的片段。
- 在486分钟真实公交视频上实现零样本乘客上下车与支付行为识别。
- 适合关注低成本、无标注数据公交智能分析的研究者与开发者。
公共交通视频理解可提供传统计数器和票务系统无法获取的细粒度数据。然而,监督视频模型需任务特定标注,直接应用视觉语言模型(VLM)于长时车载视频则不可靠且成本高。为融合两者优势,我们提出GHR-VLM:一种基于视觉定位的混合推理框架,实现零样本公交视频分析。该框架受启发于显式视觉定位能将长监控流转化为紧凑的、以乘客为中心的时空证据,从而提升VLM推理能力。具体而言,采用边缘-云协同设计:轻量级边缘监控持续追踪车门状态并分割乘客片段;后端VLM通过两阶段粗到精的时空证据优化,识别上车乘客并分类支付行为。仅对定位后的乘客片段与接触图调用VLM,显著减少云端推理开销,避免支付类别训练数据依赖,并提供VLM原本难以识别的局部化证据。在486分钟真实公交监控视频上的评估表明,该方法具备实现乘客级支付分析的潜力,但也揭示了低质量视频条件带来的挑战。
原文摘要 · Abstract (English)
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models (VLMs) directly to long onboard videos is unreliable and costly. To leverage the complementary strengths of both approaches, we propose GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics. It is motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence. Specifically, we propose an edge-cloud design in which a lightweight edge-based monitor continuously tracks door status and segments passenger clips. A backend VLM then identifies boarding passengers and classifies payment behavior through a two-stage coarse-to-fine refinement of spatiotemporal evidence. By invoking the VLM only on grounded passenger clips and contact sheets, GHR-VLM reduces cloud inference, avoids payment-specific training data, and supplies the localized evidence that VLMs otherwise struggle to identify. Evaluation on 486 minutes of real-world bus surveillance video demonstrates the potential of grounded edge-cloud reasoning for passenger-level payment analytics while highlighting the challenges posed by degraded video conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。