用语言模型透视城市认知,揭示其对典型城市的隐含偏好。
Mapping the City Through the Lens of Language Models

- 通过40项指标分析模型对城市形态的隐含假设
- 发现模型偏好面积大、发展快、基建强的城市
- 适合研究AI城市认知或模型偏见的学者
语言模型在补全城市指称时,常基于未明言的城市规模、形态、基础设施、环境与功能等假设。我们不命名具体城市,通过十组开源权重检查点,评估40个经审计指标和七个维度的匿名城市画像。方法结合约束概率评分、预设可靠性筛选、谱系感知聚合、多重人口加权、独立复现样本及全画像验证。最显著的共同趋势是偏好开发面积大、近期增长快、地图覆盖的基础设施与非住宅容量高、形态密集的城市。多数趋势在复现数据中重复出现,完整画像的直接评分与指标构建结果显示中度一致性。地理差异在控制城市规模与发展水平后减弱,可靠配对任务表明典型性与理想性常高度一致。该框架使语言模型所认为的‘普通城市’这一模糊概念变得可实证追踪,揭示出由模型决定的共通城市图景。
原文摘要 · Abstract (English)
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。