arXiv:2512.19221cs.CVcs.AI2025-12被引 1

用图像关系图提升城市感知预测准确率

From Pixels to Predicates Structuring urban perception with scene graphs

  • 将街景图转为物体-关系-物体三元组结构
  • 比纯图像模型平均提升26%预测准确率
  • 可解释哪些关系导致城市感知差

感知研究越来越多地使用街景图像,但许多方法仍依赖像素特征或物体共现统计,忽视了塑造人类感知的显式关系。本研究提出一个三阶段流程:首先使用开集全景场景图模型(OpenPSG)解析每张图像,提取物体-谓词-物体三元组;其次通过异构图自编码器(GraphMAE)学习紧凑的场景级嵌入;最后利用神经网络从这些嵌入中预测六种感知指标得分。在准确率、精确度和跨城市泛化能力上与仅使用图像的基线模型对比评估。结果表明:(i)本方法在感知预测准确率上平均提升26%;(ii)在跨城市预测任务中仍保持强泛化性能。此外,结构化表示揭示了导致城市感知评分下降的关系模式,如墙上的涂鸦、车辆停在人行道上。总体而言,该研究证明基于图的结构能提供表达性强、泛化好且可解释的信号,推动以人为本、情境感知的城市分析发展。

原文摘要 · Abstract (English)

Perception research is increasingly modelled using streetscapes, yet many approaches still rely on pixel features or object co-occurrence statistics, overlooking the explicit relations that shape human perception. This study proposes a three stage pipeline that transforms street view imagery (SVI) into structured representations for predicting six perceptual indicators. In the first stage, each image is parsed using an open-set Panoptic Scene Graph model (OpenPSG) to extract object predicate object triplets. In the second stage, compact scene-level embeddings are learned through a heterogeneous graph autoencoder (GraphMAE). In the third stage, a neural network predicts perception scores from these embeddings. We evaluate the proposed approach against image-only baselines in terms of accuracy, precision, and cross-city generalization. Results indicate that (i) our approach improves perception prediction accuracy by an average of 26% over baseline models, and (ii) maintains strong generalization performance in cross-city prediction tasks. Additionally, the structured representation clarifies which relational patterns contribute to lower perception scores in urban scenes, such as graffiti on wall and car parked on sidewalk. Overall, this study demonstrates that graph-based structure provides expressive, generalizable, and interpretable signals for modelling urban perception, advancing human-centric and context-aware urban analytics.

场景图城市感知可解释性图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。