找出影响城市视觉感知的关键视觉要素,通过可控编辑验证其作用
How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

- 用可解释的视觉杠杆框架,定位影响判断的具体局部修改
- 50个场景实验发现交通设施与维护状态对安全感知影响最大
- 适合城市规划、AI可解释性研究者参考
街景感知模型可大规模预测安全性等主观属性,但仍是相关关系,无法识别特定场景中哪些局部视觉变化可能改变人类判断。本文提出基于视觉杠杆的干预式反事实框架,将场景级可解释性转化为结构化反事实编辑的有限搜索。每个杠杆包含语义概念、空间范围、干预方向和约束编辑模板。候选编辑通过提示条件图像编辑生成,并仅保留满足同地点保持、局部性、真实性和合理性检验的结果。在来自五个城市的50个场景试点中,框架揭示了代理驱动的方向性模式,以及仅靠提示编辑时的实际失败类型,其中交通基础设施和物理维护对安全感知的辅助提升最为显著。未来验证仍需依赖人工成对判断作为基准。
原文摘要 · Abstract (English)
Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized visual changes would plausibly shift human judgement for a specific scene. We propose a lever-based interventional counterfactual framework that recasts scene-level explainability as a bounded search over structured counterfactual edits. Each lever specifies a semantic concept, spatial support, intervention direction, and constrained edit template. Candidate edits are generated through prompt-conditioned image editing and retained only if they satisfy validity checks for same-place preservation, locality, realism, and plausibility. In a pilot across 50 scenes from five cities, the framework reveals preliminary proxy-based directional patterns and a practical failure taxonomy under prompt-only editing, with Mobility Infrastructure and Physical Maintenance showing the largest auxiliary safety shifts. Human pairwise judgements remain the ground-truth endpoint for future validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。