让地图搜索理解人对城市空间的活动期待和情感偏好。
PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

- 将用户查询拆解为功能与情感子意图,结合视觉证据验证与感知偏好重排。
- 在米兰31956个街景位置上达到88.0%精度@5,优于多个基线模型。
- 适合需要理解人类体验的城市规划、导航与智能搜索系统使用。
人们搜索城市户外场所不仅关注类别或功能,还关心场所能支持的活动及其感知氛围。现有地理空间检索仍以兴趣点(POI)为中心且依赖元数据,难以满足开放性、情感性或活动导向的需求。我们提出PlaceSeek,一种以人为中心的户外场所检索框架,将自然语言查询映射到地理定位的街景图像。PlaceSeek引入意图感知检索机制,将用户查询分解为功能与情感子意图。语义对齐模块通过物理证据验证候选结果是否支持预期活动;情感对齐模块利用基于LoRA微调的视觉-语言模型(训练于人类城市感知判断数据),对物理有效的候选结果进行重排序。我们在米兰31,956个街景位置上,针对10个自然语言查询(由五名人类评估者标注)进行评估,PlaceSeek取得88.0% Precision@5、平均匹配得分3.39/4.0、0.920 nDCG@5,显著优于CLIP、微调版CLIP、SigLIP及基于VQA的基线模型。消融实验表明,物理对齐是保证检索有效性的关键,而情感对齐则提升物理有效候选者的排序质量。研究揭示复杂城市空间查询需同时建模可验证的视觉证据与人类感知偏好。PlaceSeek为下一代以人为本的地理空间检索系统提供可行框架。
原文摘要 · Abstract (English)
People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perceived. Existing geospatial retrieval remains largely POIcentric and metadata-driven, making it difficult to satisfy openended, affective, or activity-oriented needs. We present PlaceSeek, a human-centered outdoor place retrieval framework that maps natural-language queries to geolocated street-view imagery. PlaceSeek introduces an intent-aware retrieval mechanism that decomposes user queries into functional and affective sub-intents. A Semantic Grounding Module verifies whether candidate street-view results contain the physical evidence needed to support the intended activity, while an Affective Alignment Module re-ranks physically valid candidates using a LoRA-adapted vision-language model trained on human urban perception judgments. We evaluate PlaceSeek on 31,956 street-view locations in Milan across 10 naturallanguage queries annotated by five human evaluators. PlaceSeek achieves 88.0% Precision@5, a mean match score of 3.39/4.0, and 0.920 nDCG@5, outperforming CLIP, fine-tuned CLIP, SigLIP, and a VQA-based baseline. Ablation results show that physical grounding is essential for retrieval validity, while affective alignment improves ranking quality among physically valid candidates. These findings highlight that complex urban spatial queries require modeling both verifiable visual evidence and human perceptual preferences. PlaceSeek provides a potential framework for human-centered nextgeneration geospatial retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。