arXiv:2509.13484cs.CVcs.CY2025-09AAAI被引 4

让视觉模型理解城市中人群的复杂互动关系,精准定位社交群体。

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

  • 分三阶段:先检测人与深度,再用VLM判断人际关联,最后聚合定位群体。
  • 在10万张街景图上标注了个体与社交群体,覆盖真实城市场景。
  • 适合做城市规划、智能交通和人机交互研究者参考。

理解公共空间中的群体社交互动对城市规划至关重要,有助于设计更具社会活力与包容性的环境。从图像中检测此类互动需解读细微的视觉线索,如关系、距离与共同运动——这些语义复杂的信号超出了传统目标检测范畴。为此,我们提出社交群体区域检测任务,要求推断并空间定位由抽象人际关联定义的视觉区域。我们提出MINGLE(Modeling INterpersonal Group-Level Engagement),一个模块化三阶段流程:(1) 使用现成的人体检测与深度估计;(2) 基于视觉语言模型(VLM)推理成对的社会归属关系;(3) 采用轻量级空间聚合算法定位社交连通群体。为支持该任务并推动未来研究,我们构建了一个包含10万张城市街景图像的新数据集,标注了个体与社交互动群体的边界框及标签。标注结合人工标注与MINGLE输出,确保语义丰富性与现实场景广泛覆盖。

原文摘要 · Abstract (English)

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity, and co-movement - semantically complex signals that go beyond traditional object detection. To address this challenge, we introduce a social group region detection task, which requires inferring and spatially grounding visual regions defined by abstract interpersonal relations. We propose MINGLE (Modeling INterpersonal Group-Level Engagement), a modular three-stage pipeline that integrates: (1) off-the-shelf human detection and depth estimation, (2) VLM-based reasoning to classify pairwise social affiliation, and (3) a lightweight spatial aggregation algorithm to localize socially connected groups. To support this task and encourage future research, we present a new dataset of 100K urban street-view images annotated with bounding boxes and labels for both individuals and socially interacting groups. The annotations combine human-created labels and outputs from the MINGLE pipeline, ensuring semantic richness and broad coverage of real-world scenarios.

视觉语言模型群体检测城市计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。